gpt4all

AI/gpt4all

mirror of https://github.com/nomic-ai/gpt4all.git synced 2024-10-01 01:06:10 -04:00

Author	SHA1	Message	Date
Jared Van Bortel	e94177ee9a	llamamodel: fix embedding crash for >512 tokens after #2310 (#2383 ) Signed-off-by: Jared Van Bortel <jared@nomic.ai>	2024-05-29 10:51:00 -04:00
Jared Van Bortel	f047f383d0	llama.cpp: update submodule for "code" model crash workaround (#2382 ) Signed-off-by: Jared Van Bortel <jared@nomic.ai>	2024-05-29 10:50:00 -04:00
Jared Van Bortel	f1b4092ca6	llamamodel: fix BERT tokenization after llama.cpp update (#2381 ) Signed-off-by: Jared Van Bortel <jared@nomic.ai>	2024-05-28 13:11:57 -04:00
Jared Van Bortel	c779d8a32d	python: init_gpu fixes (#2368 ) * python: tweak GPU init failure message * llama.cpp: update submodule for use-after-free fix Signed-off-by: Jared Van Bortel <jared@nomic.ai>	2024-05-20 18:04:11 -04:00
Jared Van Bortel	2025d2d15b	llmodel: add CUDA to the DLL search path if CUDA_PATH is set (#2357 ) Signed-off-by: Jared Van Bortel <jared@nomic.ai>	2024-05-16 17:39:49 -04:00
Jared Van Bortel	a92d266cea	cmake: fix Metal build after #2310 (#2350 ) I don't understand why this is needed, but it works. Signed-off-by: Jared Van Bortel <jared@nomic.ai>	2024-05-15 18:12:32 -04:00
Jared Van Bortel	d2a99d9bc6	support the llama.cpp CUDA backend (#2310 ) * rebase onto llama.cpp commit ggerganov/llama.cpp@d46dbc76f * support for CUDA backend (enabled by default) * partial support for Occam's Vulkan backend (disabled by default) * partial support for HIP/ROCm backend (disabled by default) * sync llama.cpp.cmake with upstream llama.cpp CMakeLists.txt * changes to GPT4All backend, bindings, and chat UI to handle choice of llama.cpp backend (Kompute or CUDA) * ship CUDA runtime with installed version * make device selection in the UI on macOS actually do something * model whitelist: remove dbrx, mamba, persimmon, plamo; add internlm and starcoder2 Signed-off-by: Jared Van Bortel <jared@nomic.ai>	2024-05-15 15:27:50 -04:00
Jared Van Bortel	9f9d8e636f	backend: do not crash if GGUF lacks general.architecture (#2346 ) Signed-off-by: Jared Van Bortel <jared@nomic.ai>	2024-05-15 13:57:13 -04:00
Jared Van Bortel	6d8888b267	llamamodel: free the batch in embedInternal (#2348 ) Signed-off-by: Jared Van Bortel <jared@nomic.ai>	2024-05-15 12:46:12 -04:00
Jared Van Bortel	577ebd4826	mixpanel: report cpu_supports_avx2 on startup (#2299 ) Signed-off-by: Jared Van Bortel <jared@nomic.ai>	2024-05-02 16:09:41 -04:00
Jared Van Bortel	adaecb7a72	mixpanel: improved GPU device statistics (plus GPU sort order fix) (#2297 ) Signed-off-by: Jared Van Bortel <jared@nomic.ai>	2024-05-01 16:15:48 -04:00
Jared Van Bortel	c622921894	improve mixpanel usage statistics (#2238 ) Other changes: - Always display first start dialog if privacy options are unset (e.g. if the user closed GPT4All without selecting them) - LocalDocs scanQueue is now always deferred - Fix a potential crash in magic_match - LocalDocs indexing is now started after the first start dialog is dismissed so usage stats are included Signed-off-by: Jared Van Bortel <jared@nomic.ai>	2024-04-25 13:16:52 -04:00
Jared Van Bortel	ba53ab5da0	python: do not print GPU name with verbose=False, expose this info via properties (#2222 ) * llamamodel: only print device used in verbose mode Signed-off-by: Jared Van Bortel <jared@nomic.ai> * python: expose backend and device via GPT4All properties Signed-off-by: Jared Van Bortel <jared@nomic.ai> * backend: const correctness fixes Signed-off-by: Jared Van Bortel <jared@nomic.ai> * python: bump version Signed-off-by: Jared Van Bortel <jared@nomic.ai> * python: typing fixups Signed-off-by: Jared Van Bortel <jared@nomic.ai> * python: fix segfault with closed GPT4All Signed-off-by: Jared Van Bortel <jared@nomic.ai> --------- Signed-off-by: Jared Van Bortel <jared@nomic.ai>	2024-04-18 14:52:02 -04:00
Jared Van Bortel	ac498f79ac	fix regressions in system prompt handling (#2219 ) * python: fix system prompt being ignored * fix unintended whitespace after system prompt Signed-off-by: Jared Van Bortel <jared@nomic.ai>	2024-04-15 11:39:48 -04:00
Jared Van Bortel	3f8257c563	llamamodel: fix semantic typo in nomic client dynamic mode (#2216 ) Signed-off-by: Jared Van Bortel <jared@nomic.ai>	2024-04-12 17:25:15 -04:00
Jared Van Bortel	46818e466e	python: embedding cancel callback for nomic client dynamic mode (#2214 ) Signed-off-by: Jared Van Bortel <jared@nomic.ai>	2024-04-12 16:00:39 -04:00
Jared Van Bortel	459289b94c	embed4all: small fixes related to nomic client local embeddings (#2213 ) * actually submit larger batches with increased n_ctx * fix crash when llama_tokenize returns no tokens Signed-off-by: Jared Van Bortel <jared@nomic.ai>	2024-04-12 10:54:15 -04:00
Jared Van Bortel	1b84a48c47	python: add list_gpus to the GPT4All API (#2194 ) Other changes: * fix memory leak in llmodel_available_gpu_devices * drop model argument from llmodel_available_gpu_devices * breaking: make GPT4All/Embed4All arguments past model_name keyword-only Signed-off-by: Jared Van Bortel <jared@nomic.ai>	2024-04-04 14:52:13 -04:00
Jared Van Bortel	67843edc7c	backend: update llama.cpp submodule for wpm locale fix (#2163 ) Signed-off-by: Jared Van Bortel <jared@nomic.ai>	2024-03-26 11:04:22 -04:00
Jared Van Bortel	83ada4ca89	backend: update llama.cpp submodule for Unicode paths fix (#2162 ) Signed-off-by: Jared Van Bortel <jared@nomic.ai>	2024-03-26 11:01:02 -04:00
Jared Van Bortel	0455b80b7f	Embed4All: optionally count tokens, misc fixes (#2145 ) Key changes: * python: optionally return token count in Embed4All.embed * python and docs: models2.json -> models3.json * Embed4All: require explicit prefix for unknown models * llamamodel: fix shouldAddBOS for Bert and Nomic Bert Signed-off-by: Jared Van Bortel <jared@nomic.ai>	2024-03-20 11:24:02 -04:00
Jared Van Bortel	a1bb6084ed	python: documentation update and typing improvements (#2129 ) Key changes: * revert "python: tweak constructor docstrings" * docs: update python GPT4All and Embed4All documentation * breaking: require keyword args to GPT4All.generate Signed-off-by: Jared Van Bortel <jared@nomic.ai>	2024-03-19 17:25:22 -04:00
Jared Van Bortel	699410014a	fix non-AVX CPU detection (#2141 ) * chat: fix non-AVX CPU detection on Windows * bindings: throw exception instead of logging to console Signed-off-by: Jared Van Bortel <jared@nomic.ai>	2024-03-19 10:56:14 -04:00
Jared Van Bortel	255568fb9a	python: various fixes for GPT4All and Embed4All (#2130 ) Key changes: * honor empty system prompt argument * current_chat_session is now read-only and defaults to None * deprecate fallback prompt template for unknown models * fix mistakes from #2086 Signed-off-by: Jared Van Bortel <jared@nomic.ai>	2024-03-15 11:49:58 -04:00
Jared Van Bortel	53f109f519	llamamodel: fix macOS build (#2125 ) Signed-off-by: Jared Van Bortel <jared@nomic.ai>	2024-03-14 12:06:07 -04:00
Jared Van Bortel	406e88b59a	implement local Nomic Embed via llama.cpp (#2086 ) Signed-off-by: Jared Van Bortel <jared@nomic.ai>	2024-03-13 18:09:24 -04:00
Jared Van Bortel	5c248dbec9	models: new MPT model file without duplicated token_embd.weight (#2006 ) Signed-off-by: Jared Van Bortel <jared@nomic.ai>	2024-03-08 17:18:38 -05:00
Jared Van Bortel	c19b763e03	llmodel_c: expose fakeReply to the bindings (#2061 ) Signed-off-by: Jared Van Bortel <jared@nomic.ai>	2024-03-06 13:32:24 -05:00
Jared Van Bortel	f500bcf6e5	llmodel: default to a blank line between reply and next prompt (#1996 ) Also make some related adjustments to the provided Alpaca-style prompt templates and system prompts. Signed-off-by: Jared Van Bortel <jared@nomic.ai>	2024-02-26 13:11:15 -05:00
Jared Van Bortel	007d469034	bert: fix layer norm epsilon value (#1946 ) Signed-off-by: Jared Van Bortel <jared@nomic.ai>	2024-02-26 13:09:01 -05:00
Adam Treat	f720261d46	Fix another vulnerable spot for crashes. Signed-off-by: Adam Treat <treat.adam@gmail.com>	2024-02-26 12:04:16 -06:00
chrisbarrera	f8b1069a1c	add min_p sampling parameter (#2014 ) Signed-off-by: Christopher Barrera <cb@arda.tx.rr.com> Co-authored-by: Jared Van Bortel <cebtenzzre@gmail.com>	2024-02-24 17:51:34 -05:00
Jared Van Bortel	e7f2ff189f	fix some compilation warnings on macOS Signed-off-by: Jared Van Bortel <jared@nomic.ai>	2024-02-22 15:09:06 -05:00
Jared Van Bortel	88e330ef0e	llama.cpp: enable Kompute support for 10 more model arches (#2005 ) These are Baichuan, Bert and Nomic Bert, CodeShell, GPT-2, InternLM, MiniCPM, Orion, Qwen, and StarCoder. Signed-off-by: Jared Van Bortel <jared@nomic.ai>	2024-02-22 14:34:42 -05:00
Jared Van Bortel	fc6c5ea0c7	llama.cpp: gemma: allow offloading the output tensor (#1997 ) Signed-off-by: Jared Van Bortel <jared@nomic.ai>	2024-02-22 14:06:18 -05:00
Jared Van Bortel	4fc4d94be4	fix chat-style prompt templates (#1970 ) Also use a new version of Mistral OpenOrca. Signed-off-by: Jared Van Bortel <jared@nomic.ai>	2024-02-21 15:45:32 -05:00
Jared Van Bortel	7810b757c9	llamamodel: add gemma model support Signed-off-by: Jared Van Bortel <jared@nomic.ai>	2024-02-21 13:36:31 -06:00
Adam Treat	d948a4f2ee	Complete revamp of model loading to allow for more discreet control by the user of the models loading behavior. Signed-off-by: Adam Treat <treat.adam@gmail.com>	2024-02-21 10:15:20 -06:00
Jared Van Bortel	6fdec808b2	backend: update llama.cpp for faster state serialization Signed-off-by: Jared Van Bortel <jared@nomic.ai>	2024-02-13 17:39:18 -05:00
Jared Van Bortel	a1471becf3	backend: update llama.cpp for Intel GPU blacklist Signed-off-by: Jared Van Bortel <jared@nomic.ai>	2024-02-12 13:16:24 -05:00
Jared Van Bortel	eb1081d37e	cmake: fix LLAMA_DIR use before set Signed-off-by: Jared Van Bortel <jared@nomic.ai>	2024-02-09 22:00:14 -05:00
Jared Van Bortel	e60b388a2e	cmake: fix backwards LLAMA_KOMPUTE default Signed-off-by: Jared Van Bortel <jared@nomic.ai>	2024-02-09 21:53:32 -05:00
Jared Van Bortel	fc7e5f4a09	ci: fix missing Kompute support in python bindings (#1953 ) Signed-off-by: Jared Van Bortel <jared@nomic.ai>	2024-02-09 21:40:32 -05:00
Jared Van Bortel	bf493bb048	Mixtral crash fix and python bindings v2.2.0 (#1931 ) Signed-off-by: Jared Van Bortel <jared@nomic.ai>	2024-02-06 11:01:15 -05:00
Jared Van Bortel	92c025a7f6	llamamodel: add 12 new architectures for CPU inference (#1914 ) Baichuan, BLOOM, CodeShell, GPT-2, Orion, Persimmon, Phi and Phi-2, Plamo, Qwen, Qwen2, Refact, StableLM Signed-off-by: Jared Van Bortel <jared@nomic.ai>	2024-02-05 16:49:31 -05:00
Jared Van Bortel	10e3f7bbf5	Fix VRAM leak when model loading fails (#1901 ) Signed-off-by: Jared Van Bortel <jared@nomic.ai>	2024-02-01 15:45:45 -05:00
Jared Van Bortel	eadc3b8d80	backend: bump llama.cpp for VRAM leak fix when switching models Signed-off-by: Jared Van Bortel <jared@nomic.ai>	2024-01-31 17:24:01 -05:00
Jared Van Bortel	6db5307730	update llama.cpp for unhandled Vulkan OOM exception fix Signed-off-by: Jared Van Bortel <jared@nomic.ai>	2024-01-31 16:44:58 -05:00
Jared Van Bortel	0a40e71652	Maxwell/Pascal GPU support and crash fix (#1895 ) Signed-off-by: Jared Van Bortel <jared@nomic.ai>	2024-01-31 16:32:32 -05:00
Jared Van Bortel	b11c3f679e	bump llama.cpp-mainline for C++11 compat Signed-off-by: Jared Van Bortel <jared@nomic.ai>	2024-01-31 15:02:34 -05:00

1 2 3 4 5 ...

252 Commits