gpt4all

Commit Graph

Author	SHA1	Message	Date
Aaron Miller	4a24b586df	llama.cpp: metal buffer freeing	1 year ago
Aaron Miller	137bc2c367	replit: free metal context	1 year ago
Aaron Miller	57dc0c8953	adjust eval buf sizes to pass long input test	1 year ago
Aaron Miller	7a5f6e4726	limit prompt batch size to 128	1 year ago
Aaron Miller	883775bc5f	move 230511 submodule to nomic fork, fix alibi assert	1 year ago
Andriy Mulyar	46a0762bd5	Python Bindings: Improved unit tests, documentation and unification of API (#1090 ) * Makefiles, black, isort * Black and isort * unit tests and generation method * chat context provider * context does not reset * Current state * Fixup * Python bindings with unit tests * GPT4All Python Bindings: chat contexts, tests * New python bindings and backend fixes * Black and Isort * Documentation error * preserved n_predict for backwords compat with langchain --------- Co-authored-by: Adam Treat <treat.adam@gmail.com>	1 year ago
Aaron Miller	40a3faeb05	Use ggml scratch bufs for mpt and gptj models (#1104 ) * backend/gptj: use scratch buffers reduces total memory required and makes eval buf not grow with n_past * backend/mpt: use scratch bufs * fix format-related compile warnings	1 year ago
Aaron Miller	8d19ef3909	backend: factor out common elements in model code (#1089 ) * backend: factor out common structs in model code prepping to hack on these by hopefully making there be fewer places to fix the same bug rename * use common buffer wrapper instead of manual malloc * fix replit compile warnings	1 year ago
Aaron Miller	28d41d4f6d	falcon: use model-local eval & scratch bufs (#1079 ) fixes memory leaks copied from ggml/examples based implementation	1 year ago
Zach Nussbaum	2565f6a94a	feat: add conversion script	1 year ago
Aaron Miller	198b5e4832	add Falcon 7B model Tested with https://huggingface.co/TheBloke/falcon-7b-instruct-GGML/blob/main/falcon7b-instruct.ggmlv3.q4_0.bin	1 year ago
Aaron Miller	db34a2f670	llmodel: skip attempting Metal if model+kvcache > 53% of system ram	1 year ago
Aaron Miller	b19a3e5b2c	add requiredMem method to llmodel impls most of these can just shortcut out of the model loading logic llama is a bit worse to deal with because we submodule it so I have to at least parse the hparams, and then I just use the size on disk as an estimate for the mem size (which seems reasonable since we mmap() the llama files anyway)	1 year ago
Adam Treat	a0f80453e5	Use sysinfo in backend.	1 year ago
niansa/tuxifan	47323f8591	Update replit.cpp replit_tokenizer_detokenize returnins std::string now Signed-off-by: niansa/tuxifan <tuxifan@posteo.de>	1 year ago
niansa	0855c0df1d	Fixed Replit implementation compile warnings	1 year ago
Aaron Miller	1290b32451	update to latest mainline llama.cpp add max_size param to ggml_metal_add_buffer - introduced in https://github.com/ggerganov/llama.cpp/pull/1826	1 year ago
niansa/tuxifan	5eee16c97c	Do not specify "success" as error for unsupported models Signed-off-by: niansa/tuxifan <tuxifan@posteo.de>	1 year ago
Adam Treat	bd58c46da0	Initialize these to nullptr to prevent double deletion when a model fails to load.	1 year ago
niansa/tuxifan	68f9786ed9	Use operator ""_MiB (#991 )	1 year ago
Aaron Miller	abc081e48d	fix llama.cpp k-quants (#988 ) * enable k-quants on all mainline builds	1 year ago
Aaron Miller	c4319d2c8e	dlhandle: prevent libs from using each other's symbols (#977 ) use RTLD_LOCAL so that symbols are only exposed via dlsym without this all symbols exported by the libs are available for symbol resolution, resulting in different lib versions potentially resolving each other's symbols, causing incredibly cursed behavior such as https://gist.github.com/apage43/085c1ff69f6dd05387793ebc301840f6	1 year ago
Aaron Miller	f71d8efc71	metal replit (#931 ) metal+replit makes replit work with Metal and removes its use of `mem_per_token` in favor of fixed size scratch buffers (closer to llama.cpp)	1 year ago
Aaron Miller	85964a7635	bump llama.cpp mainline to latest (#964 )	1 year ago
Tim Miller	797891c995	Initial Library Loader for .NET Bindings / Update bindings to support newest changes (#763 ) * Initial Library Loader * Load library as part of Model factory * Dynamically search and find the dlls * Update tests to use locally built runtimes * Fix dylib loading, add macos runtime support for sample/tests * Bypass automatic loading by default. * Only set CMAKE_OSX_ARCHITECTURES if not already set, allow cross-compile * Switch Loading again * Update build scripts for mac/linux * Update bindings to support newest breaking changes * Fix build * Use llmodel for Windows * Actually, it does need to be libllmodel * Name * Remove TFMs, bypass loading by default * Fix script * Delete mac script --------- Co-authored-by: Tim Miller <innerlogic4321@ghmail.com>	1 year ago
Aaron Miller	88616fde7f	llmodel: change tokenToString to not use string_view (#968 ) fixes a definite use-after-free and likely avoids some other potential ones - std::string will convert to a std::string_view automatically but as soon as the std::string in question goes out of scope it is already freed and the string_view is pointing at freed memory - this is mostly fine if its returning a reference to the tokenizer's internal vocab table but it's, imo, too easy to return a reference to a dynamically constructed string with this as replit is doing (and unfortunately needs to do to convert the internal whitespace replacement symbol back to a space)	1 year ago
Adam Treat	84deebd223	Fix compile for windows and linux again. PLEASE DON'T REVERT THISgit gui!	1 year ago
Juuso Alasuutari	5cfb1bda89	llmodel: add model wrapper destructor, fix mem leak in golang bindings (#862 ) Signed-off-by: Juuso Alasuutari <juuso.alasuutari@gmail.com>	1 year ago
Cosmic Snow	ae4a275bcd	Fix Windows MSVC AVX builds - bug introduced in `0cb2b86730` - currently getting: `warning C5102: ignoring invalid command-line macro definition '/arch:AVX2'` - solution is to use `_options(...)` not `_definitions(...)`	1 year ago
Adam Treat	b906fb4057	When recalculating context we can't erase the BOS.	1 year ago
Aaron Miller	d3ba1295a7	Metal+LLama take two (#929 ) Support latest llama with Metal --------- Co-authored-by: Adam Treat <adam@nomic.ai> Co-authored-by: niansa/tuxifan <tuxifan@posteo.de>	1 year ago
Adam Treat	b162b5c64e	Revert "llama on Metal (#885 )" This reverts commit `c55f81b860`.	1 year ago
Aaron Miller	c55f81b860	llama on Metal (#885 ) Support latest llama with Metal --------- Co-authored-by: Adam Treat <adam@nomic.ai> Co-authored-by: niansa/tuxifan <tuxifan@posteo.de>	1 year ago
niansa/tuxifan	14e9ccbc6a	Do auto detection by default in C++ API Signed-off-by: niansa/tuxifan <tuxifan@posteo.de>	1 year ago
niansa/tuxifan	f03da8d732	Removed double-static from variables in replit.cpp The anonymous namespace already makes it static. Signed-off-by: niansa/tuxifan <tuxifan@posteo.de>	1 year ago
niansa	0cb2b86730	Synced llama.cpp.cmake with upstream	1 year ago
Aaron Miller	47fbc0e309	non-llama: explicitly greedy sampling for temp<=0 (#901 ) copied directly from llama.cpp - without this temp=0.0 will just scale all the logits to infinity and give bad output	1 year ago
Aaron Miller	b14953e136	sampling: remove incorrect offset for n_vocab (#900 ) no effect, but avoids a potential bug later if we use actualVocabSize - which is for when a model has a larger embedding tensor/# of output logits than actually trained token to allow room for adding extras in finetuning - presently all of our models have had "placeholder" tokens in the vocab so this hasn't broken anything, but if the sizes did differ we want the equivalent of `logits[actualVocabSize:]` (the start point is unchanged), not `logits[-actualVocabSize:]` (this.)	1 year ago
Adam Treat	010a04d96f	Revert "Synced llama.cpp.cmake with upstream (#887 )" This reverts commit `89910c7ca8`.	1 year ago
Adam Treat	7e304106cc	Fix for windows.	1 year ago
niansa/tuxifan	89910c7ca8	Synced llama.cpp.cmake with upstream (#887 )	1 year ago
Richard Guo	c4706d0c14	Replit Model (#713 ) * porting over replit code model to gpt4all * replaced memory with kv_self struct * continuing debug * welp it built but lot of sus things * working model loading and somewhat working generate.. need to format response? * revert back to semi working version * finally got rid of weird formatting * figured out problem is with python bindings - this is good to go for testing * addressing PR feedback * output refactor * fixed prompt reponse collection * cleanup * addressing PR comments * building replit backend with new ggmlver code * chatllm replit and clean python files * cleanup * updated replit to match new llmodel api * match llmodel api and change size_t to Token * resolve PR comments * replit model commit comment	1 year ago
Adam Treat	c5de9634c9	Fix llama models on linux and windows.	1 year ago
Adam Treat	8a9ad258f4	Fix symbol resolution on windows.	1 year ago
Adam Treat	812b2f4b29	Make installers work with mac/windows for big backend change.	1 year ago
Adam Treat	f73333c6a1	Update to latest llama.cpp	1 year ago
Adam Treat	301d2fdbea	Fix up for newer models on reset context. This fixes the model from totally failing after a reset context.	1 year ago
AT	5f95aa9fc6	We no longer have an avx_only repository and better error handling for minimum hardware requirements. (#833 )	1 year ago
AT	bbe195ee02	Backend prompt dedup (#822 ) * Deduplicated prompt() function code	1 year ago
Ikko Eltociear Ashimine	945297d837	Update README.md huggingface -> Hugging Face Signed-off-by: Ikko Eltociear Ashimine <eltociear@gmail.com>	1 year ago

1 2

92 Commits (6d9575e1033d4745bfe85d958086ec0549f8e157)