gpt4all/gpt4all-backend/scripts
Aaron Miller bbcee1ced5 New tokenizer implementation for MPT and GPT-J
Improves output quality by making these tokenizers more closely
match the behavior of the huggingface `tokenizers` based BPE
tokenizers these models were trained with.

Featuring:
 * Fixed unicode handling (via ICU)
 * Fixed BPE token merge handling
 * Complete added vocabulary handling
2023-05-30 12:05:57 -04:00
..
convert_mpt_hf_to_ggml.py fix: use right conversion script 2023-05-11 11:20:43 -04:00
gen_tokenizer_include.py New tokenizer implementation for MPT and GPT-J 2023-05-30 12:05:57 -04:00