petals

mirror of https://github.com/bigscience-workshop/petals synced 2024-10-31 09:20:41 +00:00

Author	SHA1	Message	Date
Alexander Borzunov	523a7cad33	Fix issues related to `petals` as a module (#159 ) 1. Added `from petals.client import *` to `petals/__init__.py`, so you can write just that: ```python from petals import DistributedBloomForCausalLM ``` I didn't do the same with server, since its classes are supposed to by used by `petals.cli.run_server`, not end-users. Though it's still possible to do `from petals.server.smth import smth` if necessary. 2. Fixed one more logging issue: log lines from hivemind were shown twice due to a bug in #156. 3. Removed unused `runtime.py`, since the server actually uses `hivemind.moe.Runtime`, and `runtime.py` has no significant changes comparing to it.	2022-12-16 09:09:06 +04:00
justheuristic	91898c3c90	Switch to speedtest-cli (#157 ) This pullrequest removes custom speed_test code in favour of speedtest-cli module. This is necessary to ensure that random warnings / print-outs do not mess with our outputs. Co-authored-by: Max Ryabinin <mryabinin0@gmail.com>	2022-12-15 15:21:33 +03:00
Max Ryabinin	34644f13e1	Downgrade CUDA in Docker image to 11.0.3 (#145 ) * Downgrade CUDA in Docker image to 11.0.3 * Remove development deps from the image	2022-12-15 10:14:29 +03:00
Artem Chumachenko	7911c2641d	Update advanced notebooks (#148 ) Update examples Co-authored-by: Alexander Borzunov <borzunov.alexander@gmail.com>	2022-12-15 09:28:25 +04:00
Alexander Borzunov	668b736031	Fix logging: do not duplicate lines, enable colors in Colab (#156 )	2022-12-15 09:12:18 +04:00
Alexander Borzunov	041ad20891	Check reachability automatically and give advice how to fix it (#155 ) 1. If we connect to the public swarm, the server now automatically checks its DHT's reachability from the outside world using API at http://health.petals.ml This is important to disallow unreachable servers to proceed (they create issues for the clients, such as repetitive retries). If http://health.petals.ml is down, the server proceeds without the check (so we don't depend on it). However, if health.petals.ml is up and explicitly tells us that we are unrechable, the server shows the reason of that and how to solve it. The check may be disabled with the `--skip_reachability_check` option (though I can't imagine cases where someone needs to use it). 2. Added `--port` and `--public_ip` as simplified convenience options for users not familiar with `--host_maddrs` and `--announce_maddrs`.	2022-12-15 05:04:09 +04:00
Alexander Borzunov	73df69a117	Reset MemoryCache during rebalancings (#154 ) Before this PR, if there were open inference sessions right when rebalancing is triggered, their cache was never properly destroyed.	2022-12-15 00:11:46 +04:00
Max Ryabinin	bd91be27ea	Add missing methods for SamplingAlgorithm, fix docstrings (#107 ) * Add missing methods for SamplingAlgorithm, fix docstrings * Add SamplingAlgorithm to _choose_sample_algorithm * Add test_sampling * Add a warning if sampling options were passed, but do_sample=False * Skip the sampling test for now Co-authored-by: Alexander Borzunov <borzunov.alexander@gmail.com>	2022-12-13 20:09:15 +03:00
Max Ryabinin	a0e8bbd28d	Fix arguments in remove_old_models.py (#153 ) * Fix arguments in remove_old_models.py * Remove unnecessary args.author * Fix the GitHub Action as well	2022-12-13 19:01:12 +03:00
Alexander Borzunov	701ec7e53e	Clean up disk space (#152 )	2022-12-13 18:50:43 +04:00
justheuristic	b04982c1a2	Bump transformers to 4.25.1 (#151 ) - latest accelerate, transformers, huggingface_hub - rearrange attention caches to support https://github.com/huggingface/transformers/pull/18344 - remove unused code - fix edge case where session crashes when receiving seq length 0 - assert transformer version when importing WrappedBloomBlock Co-authored-by: Alexander Borzunov <borzunov.alexander@gmail.com> Co-authored-by: Max Ryabinin <mryabinin0@gmail.com>	2022-12-13 11:03:49 +03:00
Alexander Borzunov	e4dc938dfe	Fix OOMs during server rebalancing (#150 ) The cause of OOMs were the cyclic references `TransformerBackend <-> PrioritizedTaskPool` that could not have been garbage collected properly. Still, I've added explicit tensor removal just in case.	2022-12-13 00:44:40 +04:00
Alexander Borzunov	83d9493b6c	Improve block size calculations (#149 )	2022-12-12 13:15:23 +04:00
Aleksandr Borzunov	f42e559c77	Update README.md	2022-12-09 17:54:00 +00:00
Alexander Borzunov	6beb686909	Add link to privacy & security Wiki (#144 )	2022-12-09 21:40:29 +04:00
Alexander Borzunov	84fec81543	Suppress asyncio error logs by default (#142 )	2022-12-09 04:32:26 +04:00
Alexander Borzunov	e99bf36647	Use common folder for all caches, make it a volume in Dockerfile (#141 )	2022-12-09 03:54:57 +04:00
Alexander Borzunov	5f50ea9c79	Update Anaconda instructions (#140 )	2022-12-09 01:16:33 +04:00
Alexander Borzunov	e1d8793f00	Show route on client (#139 )	2022-12-08 20:26:33 +04:00
Alexander Borzunov	4cb0ac4718	Update texts in "Terms of use" and "Privacy and security" sections (#138 )	2022-12-08 04:41:02 +04:00
Alexander Borzunov	a94c91d870	Add Docker commands, use permanent Discord links (#137 )	2022-12-08 01:54:22 +04:00
Alexander Borzunov	77a00e17f0	Fix "could not unlink the shared memory file" during rebalancing (#135 )	2022-12-07 12:59:34 +04:00
Alexander Borzunov	318d690a5c	Fix waiting until free memory is available (#136 )	2022-12-07 02:29:54 +04:00
Alexander Borzunov	e8fac92e59	Allow .generate() to reuse existing inference session (#132 )	2022-12-06 00:20:26 +04:00
Alexander Borzunov	1fe3716589	Don't ban servers in case of client-caused handler errors (#134 )	2022-12-05 19:05:43 +04:00
Alexander Borzunov	66f1799d32	Set default --step_timeout to 5 min (#133 )	2022-12-05 13:44:18 +04:00
Alexander Borzunov	b873d92ffa	Update README.md	2022-12-04 22:48:51 +04:00
Alexander Borzunov	5d5d2666b8	Mention parallel inference	2022-12-04 22:48:32 +04:00
Alexander Borzunov	955eae30b3	Mention 1 sec/token explicitly	2022-12-04 22:10:15 +04:00
Alexander Borzunov	33c210b973	Update Colab notebook	2022-12-03 20:38:18 +04:00
Alexander Borzunov	f56edaa13f	Fix inference and rpc_info() fault tolerance (#131 )	2022-12-03 19:28:15 +04:00
justheuristic	79a4308992	Clear trigger before engaging in update (#130 ) Update sequence_manager.py	2022-12-03 17:42:52 +03:00
Alexander Borzunov	b8e1c1b7f5	Revert to hivemind==1.1.3 for stability (#129 )	2022-12-03 17:36:05 +04:00
justheuristic	68c85e7492	Avoid synchronous updates, ban peers based on request outcome (#127 ) - sequence_manager now takes care for its own updated-ness - no need to manually update it - if a peer fails a request, sequence manager will ban this peer temporarily. Ban times increase with failure streaks Co-authored-by: Alexander Borzunov <borzunov.alexander@gmail.com>	2022-12-03 16:13:15 +03:00
Alexander Borzunov	9dbf5e2e6f	Set dht.num_workers = n_layer, update_period = 150, expiration = 300 (#125 )	2022-12-03 15:26:57 +04:00
Max Ryabinin	3ca8b4f082	Fix typos with codespell (#126 )	2022-12-03 14:19:37 +03:00
justheuristic	8491ed2bd3	Add checks for forward() inputs on the client side (#123 )	2022-12-03 15:02:48 +04:00
Max Ryabinin	055f85b83e	Call block.load_state_dict only once (#124 )	2022-12-03 15:01:56 +04:00
Artem Chumachenko	0855aa7347	Update notebooks to use full BLOOM-176B (#104 ) Co-authored-by: Alexander Borzunov <borzunov.alexander@gmail.com>	2022-12-03 14:09:21 +04:00
Max Ryabinin	4ffb4d83c7	Remove "-r" when installing Petals in examples (#122 )	2022-12-03 11:21:45 +04:00
Alexander Borzunov	d29ef70c85	Update README.md	2022-12-03 01:14:02 +04:00
Alexander Borzunov	1d9aa77697	Update README.md	2022-12-03 00:46:52 +04:00
Alexander Borzunov	da36470a4b	Update README.md	2022-12-03 00:46:08 +04:00
Alexander Borzunov	81b94df14b	Rework readme, move code example to the top, link draft of Colab (#118 )	2022-12-03 00:17:57 +04:00
Alexander Borzunov	893987ebf8	Require hivemind==1.1.4 with p2pd v0.3.13 (#121 )	2022-12-03 00:16:14 +04:00
Alexander Borzunov	fc6722576b	Choose --num_blocks for bigscience/bloom-petals automatically (#119 )	2022-12-02 23:17:44 +04:00
Alexander Borzunov	f72c220404	Suppress quantization warning and fix dtype defaults in compute benchmark (#117 )	2022-12-02 20:07:28 +04:00
Alexander Borzunov	643a054170	Make server use smart defaults (#115 ) Summary: ```python parser.add_argument('--attn_cache_size', type=str, default=None, help='The size of GPU memory allocated for storing past attention keys/values between inference steps. ' 'Examples: 500MB, 1.2GB, 1073741824 (bytes). Note that 1KB != 1KiB here. ' 'Default: 0.5GiB * num_blocks * hidden_size / 14336. ' 'The latter is the hidden size of the bigscience/bloom-petals model.') parser.add_argument('--request_timeout', type=float, required=False, default=3 * 60, help='Timeout (in seconds) for the whole rpc_forward/rpc_backward/rpc_forward_stream/rpc_backward_stream request') parser.add_argument('--session_timeout', type=float, required=False, default=30 * 60, help='Timeout (in seconds) for the whole inference session') parser.add_argument('--step_timeout', type=float, required=False, default=60, help="Timeout (in seconds) for waiting the next step's inputs inside an inference session") parser.add_argument('--load_in_8bit', type=bool, default=None, help="Convert the loaded model into mixed-8bit quantized model. Default: True if GPU is available") ``` Co-authored-by: justheuristic <justheuristic@gmail.com>	2022-12-02 17:36:39 +04:00
justheuristic	9e11f73242	Fix tile size on ampere (#116 ) Fix tile size on ampere Co-authored-by: Aleksandr Borzunov <borzunov.alexander@gmail.com>	2022-12-02 16:16:42 +03:00
justheuristic	617d70f7dc	Support --load_in_8bit on pre-Turing GPUs (#113 ) - Linear8bitLt now supports for pre-turing GPUs by temporarily upcasting quantized weights. - added a test for linear8bitlt accuracy with the new fallback, the accuracy is similar than the real thing, (slightly better due to non-quantized A) - performance is roughly halfway between the default mode and memory_efficient_backward Alternatives considered: - cupy - slow, casting to float internally - triton - fast but unstable af. every 3rd attempt to matmul is a segfault - bnb.functional.igemm (no lt) - "CuBLAS Error 8" on old GPUs Co-authored-by: Aleksandr Borzunov <borzunov.alexander@gmail.com>	2022-12-02 15:10:24 +03:00

... 3 4 5 6 7 ...

488 Commits