petals

Commit Graph

Author	SHA1	Message	Date
Alexander Borzunov	712f5a330f	Remove backup bootstrap peer	1 year ago
justheuristic	d1fa5eb260	hotfix: add initial peer that did not crash :) (#181 ) add hotfix initial peer (@borzunov's peers are down)	1 year ago
Alexander Borzunov	6dd9a938bd	Import bitsandbytes only if it's going to be used (#180 )	1 year ago
Alexander Borzunov	e27706358c	Use slightly less memory in .generate() (#177 )	1 year ago
Alexander Borzunov	55698381d0	Disable chunked_forward() on AVX512 CPUs (#179 )	1 year ago
Alexander Borzunov	6948a0c5ee	Allow to disable chunked forward (#176 )	1 year ago
Alexander Borzunov	356e099c3d	Make Docker command more visible (#175 )	1 year ago
justheuristic	ae9e71fe8e	Add local tensor-parallel fwd/bwd (#143 ) This pull request adds an option to run Petals server on multiple local GPUs. It uses https://github.com/BlackSamorez/tensor_parallel - 8bit approximation error same as in main (mean~=2% q0.9~=5%) - TP=1, 2, 3 (see screenshots above) - forward, grad w.r.t. input and inference exact match with main with TP=1 - `>=`80% GPU utilization with 3x 1080ti, batch = 8 tokens - throughput measured with and without TP - TP on 1080Tis has near-linear speedup comparable to the benchmarks (see first message) Co-authored-by: Iaroslav Lisniak <yalisnyak@nes.ru> Co-authored-by: Andrei Panferov <andrei@blacksamorez.ru> Co-authored-by: Alexander Borzunov <borzunov.alexander@gmail.com>	1 year ago
Alexander Borzunov	779959bc70	Add link to PyPI (#173 )	1 year ago
Alexander Borzunov	cdc3b6a25a	Add PyPI badge, update instructions and links in readme (#172 )	1 year ago
Aleksandr Borzunov	ff8ade8d3b	Bump version to 1.0.0	1 year ago
justheuristic	4014442a0f	Fix instruction for developers (#170 )	1 year ago
Alexander Borzunov	26e6120288	Fix code example in readme (#169 ) Makes it closer to runnable code, except for imports and defining tokenizer & data loader.	1 year ago
Alexander Borzunov	0b0277ed6f	Add link to chat.petals.ml (#168 )	1 year ago
Vadim Peretokin	50fb8205de	Correct grammar in readme (#166 )	1 year ago
Alexander Borzunov	714da529e6	Update wording in readme (#165 )	1 year ago
Alexander Borzunov	9997ada3bb	Shield alloc & free from cancellation (#163 ) A handler's RPC code may be cancelled due to a request timeout or a client closing the connection. Before this PR: - If `.cancel()` happens while waiting for `hivemind.utils.enter_asynchronously()`, the lock will never be released. - If `.cancel()` happens while doing that before freeing memory, the memory will never be freed. This PR fixes it by deferring the cancellation with [asyncio.shield()](https://docs.python.org/3/library/asyncio-task.html#asyncio.shield). Now, the cancellation will happen only when all locks are released and alloc/free has completed.	1 year ago
Alexander Borzunov	d6992fca63	Hot fix: Increase hivemind.P2P's startup_timeout for Colab, remove absent initial peer (#162 )	1 year ago
Artem Chumachenko	0a6b5f31aa	Fix misstypos in the example notebooks. (#161 ) Fix misstypos	1 year ago
Alexander Borzunov	7cdc57a04b	Alloc inference cache as one contiguous buffer (#160 )	1 year ago
Alexander Borzunov	523a7cad33	Fix issues related to `petals` as a module (#159 ) 1. Added `from petals.client import *` to `petals/__init__.py`, so you can write just that: ```python from petals import DistributedBloomForCausalLM ``` I didn't do the same with server, since its classes are supposed to by used by `petals.cli.run_server`, not end-users. Though it's still possible to do `from petals.server.smth import smth` if necessary. 2. Fixed one more logging issue: log lines from hivemind were shown twice due to a bug in #156. 3. Removed unused `runtime.py`, since the server actually uses `hivemind.moe.Runtime`, and `runtime.py` has no significant changes comparing to it.	1 year ago
justheuristic	91898c3c90	Switch to speedtest-cli (#157 ) This pullrequest removes custom speed_test code in favour of speedtest-cli module. This is necessary to ensure that random warnings / print-outs do not mess with our outputs. Co-authored-by: Max Ryabinin <mryabinin0@gmail.com>	1 year ago
Max Ryabinin	34644f13e1	Downgrade CUDA in Docker image to 11.0.3 (#145 ) * Downgrade CUDA in Docker image to 11.0.3 * Remove development deps from the image	1 year ago
Artem Chumachenko	7911c2641d	Update advanced notebooks (#148 ) Update examples Co-authored-by: Alexander Borzunov <borzunov.alexander@gmail.com>	1 year ago
Alexander Borzunov	668b736031	Fix logging: do not duplicate lines, enable colors in Colab (#156 )	1 year ago
Alexander Borzunov	041ad20891	Check reachability automatically and give advice how to fix it (#155 ) 1. If we connect to the public swarm, the server now automatically checks its DHT's reachability from the outside world using API at http://health.petals.ml This is important to disallow unreachable servers to proceed (they create issues for the clients, such as repetitive retries). If http://health.petals.ml is down, the server proceeds without the check (so we don't depend on it). However, if health.petals.ml is up and explicitly tells us that we are unrechable, the server shows the reason of that and how to solve it. The check may be disabled with the `--skip_reachability_check` option (though I can't imagine cases where someone needs to use it). 2. Added `--port` and `--public_ip` as simplified convenience options for users not familiar with `--host_maddrs` and `--announce_maddrs`.	1 year ago
Alexander Borzunov	73df69a117	Reset MemoryCache during rebalancings (#154 ) Before this PR, if there were open inference sessions right when rebalancing is triggered, their cache was never properly destroyed.	1 year ago
Max Ryabinin	bd91be27ea	Add missing methods for SamplingAlgorithm, fix docstrings (#107 ) * Add missing methods for SamplingAlgorithm, fix docstrings * Add SamplingAlgorithm to _choose_sample_algorithm * Add test_sampling * Add a warning if sampling options were passed, but do_sample=False * Skip the sampling test for now Co-authored-by: Alexander Borzunov <borzunov.alexander@gmail.com>	1 year ago
Max Ryabinin	a0e8bbd28d	Fix arguments in remove_old_models.py (#153 ) * Fix arguments in remove_old_models.py * Remove unnecessary args.author * Fix the GitHub Action as well	1 year ago
Alexander Borzunov	701ec7e53e	Clean up disk space (#152 )	1 year ago
justheuristic	b04982c1a2	Bump transformers to 4.25.1 (#151 ) - latest accelerate, transformers, huggingface_hub - rearrange attention caches to support https://github.com/huggingface/transformers/pull/18344 - remove unused code - fix edge case where session crashes when receiving seq length 0 - assert transformer version when importing WrappedBloomBlock Co-authored-by: Alexander Borzunov <borzunov.alexander@gmail.com> Co-authored-by: Max Ryabinin <mryabinin0@gmail.com>	1 year ago
Alexander Borzunov	e4dc938dfe	Fix OOMs during server rebalancing (#150 ) The cause of OOMs were the cyclic references `TransformerBackend <-> PrioritizedTaskPool` that could not have been garbage collected properly. Still, I've added explicit tensor removal just in case.	1 year ago
Alexander Borzunov	83d9493b6c	Improve block size calculations (#149 )	1 year ago
Aleksandr Borzunov	f42e559c77	Update README.md	1 year ago
Alexander Borzunov	6beb686909	Add link to privacy & security Wiki (#144 )	1 year ago
Alexander Borzunov	84fec81543	Suppress asyncio error logs by default (#142 )	1 year ago
Alexander Borzunov	e99bf36647	Use common folder for all caches, make it a volume in Dockerfile (#141 )	1 year ago
Alexander Borzunov	5f50ea9c79	Update Anaconda instructions (#140 )	1 year ago
Alexander Borzunov	e1d8793f00	Show route on client (#139 )	1 year ago
Alexander Borzunov	4cb0ac4718	Update texts in "Terms of use" and "Privacy and security" sections (#138 )	1 year ago
Alexander Borzunov	a94c91d870	Add Docker commands, use permanent Discord links (#137 )	1 year ago
Alexander Borzunov	77a00e17f0	Fix "could not unlink the shared memory file" during rebalancing (#135 )	1 year ago
Alexander Borzunov	318d690a5c	Fix waiting until free memory is available (#136 )	1 year ago
Alexander Borzunov	e8fac92e59	Allow .generate() to reuse existing inference session (#132 )	1 year ago
Alexander Borzunov	1fe3716589	Don't ban servers in case of client-caused handler errors (#134 )	1 year ago
Alexander Borzunov	66f1799d32	Set default --step_timeout to 5 min (#133 )	1 year ago
Alexander Borzunov	b873d92ffa	Update README.md	1 year ago
Alexander Borzunov	5d5d2666b8	Mention parallel inference	1 year ago
Alexander Borzunov	955eae30b3	Mention 1 sec/token explicitly	1 year ago
Alexander Borzunov	33c210b973	Update Colab notebook	1 year ago

... 3 4 5 6 7 ...

508 Commits (main) All Branches Search

508 Commits (main)

All Branches