petals

Commit Graph

Author	SHA1	Message	Date
Alexander Borzunov	5ce4f1a159	Store (start_block, end_block) in each DHT record for reliability (#510 ) This PR fixes gaps in the DHT server info caused by unavailable DHT keys. Now, one DHT key is enough to get info about all blocks hosted by a server - so we'll see info until all keys are unavailable. Also, this PR refactors `petals.client.routing` and `petals.server.block_selection` modules to use the common `compute_spans()` function (defined in `petals.utils.dht`) and `RemoteSpanInfo` class (defined in `petals.data_structures`).	8 months ago
Alexander Borzunov	dd4a3230bc	Add Falcon support (#499 ) This PR adds: - Support for models based on `transformers.FalconModel` (the in-library format for Falcon). Tested on Falcon-40B. - CI tests for Falcon-RW-1B. - `--throughput dry_run` option to evaluate throughput and exit right away (implemented by @mryab). Limitations: - Backward pass support is broken for now, will be fixed in #500. Co-authored-by: Max Ryabinin <mryabinin0@gmail.com>	9 months ago
Alexander Borzunov	6ef6bf5fa2	Create model index in DHT (#491 ) This PR creates an index of models hosted in the swarm - it is useful to know which custom models users run and display them at https://health.petals.dev as "not officially supported" models.	9 months ago
Alexander Borzunov	dc0072fde1	Wait for DHT storing state OFFLINE on shutdown (#486 )	9 months ago
Alexander Borzunov	26ebbfe8f0	Support macOS (#477 ) This PR makes both clients and servers work on macOS. Specifically, it: - Follows https://github.com/learning-at-home/hivemind/pull/586 to run a macOS-compatible `p2pd` binary (both x86-64 and ARM64 are supported) - Fixes forking issues and tests on macOS, Python 3.10+ - Introduces basic support for serving model blocks on Apple M1/M2 GPUs (torch.mps) - Increases max number of open files by default (it's not enough on Linux and is really small on macOS)	9 months ago
justheuristic	c08d09c4d3	Rewrite MemoryCache alloc_timeout logic (#434 ) - rpc_inference: server will now accept allocation timeout from user, defaults to no timeout - bugfix: inference timeout is now measured from the moment the request is received - previously, you would have to wait for your timeout plus the time it takes to sort through the queue (other users' timeout) - now, you get AllocationFailed if you had to wait for over (timeout) seconds - regardless of other users - a request for inference with no timeout will now fail instantly if there is not enough memory available - dtype number of bytes is now correctly determined for int, bool & other types --------- Co-authored-by: Your Name <you@example.com> Co-authored-by: Alexander Borzunov <borzunov.alexander@gmail.com> Co-authored-by: Aleksandr Borzunov <hxrussia@gmail.com>	9 months ago
Alexander Borzunov	063e94b4c8	Move SequenceManagerConfig -> ClientConfig, petals.dht_utils -> petals.utils.dht (#463 )	9 months ago
Alexander Borzunov	056f22515a	Prioritize short inference, unmerge pools for long inference (#458 ) Right now, long inference requests may occupy Runtime for a few seconds without giving it away to process short (most latency-sensitive requests). This PR fixes it by disallowing the merged pool for long requests and prioritizing the short ones.	9 months ago
Alexander Borzunov	8c546d988a	Test Llama, rebalancing, throughput eval, and all CLI scripts (#452 ) This PR extends CI to: 1. Test Llama code using [TinyLlama-v0](https://huggingface.co/Maykeye/TinyLLama-v0). 2. Test rebalancing (sets up a situation where the 1st server needs to change its original position). 3. Check if benchmark scripts run (in case someone breaks its code). Note that the benchmark results are meaningless here (since they're measured on a tiny swarm of CPU servers, with low `--n_steps`). 4. Test `petals.cli.run_dht`. 5. Increase swap space and watch free RAM (a common issue is that actions are cancelled without explanation if there's not enough RAM - so it's a useful reminder + debug tool). 6. Fix flapping tests for bloom-560m by increasing tolerance. Other minor changes: fix `--help` messages to show defaults, fix docs, tune rebalancing constants.	10 months ago
Vadim Peretokin	d0b5af34cd	Fix typo and make blocks message more informative (#437 ) The message really doesn't tell me much as a user, since I never touched update_period to begin with: ``` Aug 06 09:43:07.287 [WARN] [petals.server.server.run:701] Declaring blocs to DHT takes more than --update_period, consider increasing it ``` Made it better and more informative.	10 months ago
Alexander Borzunov	351e96bc46	Penalize servers that use relays during rebalancing (#428 ) Servers accessible only via relays may introduce issues if they are the only type of servers holding certain blocks. Specifically, a connection to such servers may be unstable or opened after a certain delay. This PR changes their self-reported throughput, so that the rebalancing algorithm prefers to put directly available servers for hosting each block.	10 months ago
Alexander Borzunov	fd19c21859	Update --update_period and --expiration defaults (#410 )	10 months ago
justheuristic	5af04524dd	Split long sequences into chunks (#403 ) This PR is designed to avoid OOMs when processing long sequences that happen due to the huge attention logits matrices. Co-authored-by: Alexander Borzunov <borzunov.alexander@gmail.com>	10 months ago
Alexander Borzunov	8666653cf5	Fix routing through relay, default network RPS, --token, logging, readme (#399 ) * Hide GeneratorExit in _iterate_inference_steps() * Update README.md about `--public_name` * Use .from_pretrained(..., use_auth_token=token) instead of token=token until it's fully supported across HF libs * Use default network speed 25 Mbit/s * Apply relay penalty in max-throughput routing * Replace RPS with "tokens/sec per block" in logs * Increase default expiration	10 months ago
Alexander Borzunov	b6b3ae964f	Fix --attn_cache_tokens default (#392 )	10 months ago
Alexander Borzunov	057a2fb5de	Support Llama 2 (#379 )	10 months ago
Alexander Borzunov	3218534745	Fix --token arg (#378 )	10 months ago
justheuristic	5a8de2f1f8	Fix handler memory leak, get rid of mp.Manager (#373 ) This PR removes the memory leak from somewhere within handler.py that has something to do with mp.SyncManager.	10 months ago
Alexander Borzunov	c735dd7ba3	Update transformers to 4.31.0 and peft to 0.4.0 (#371 )	10 months ago
Alexander Borzunov	a6fdfc0556	Fix AssertionError on rebalancing (#370 )	10 months ago
Alexander Borzunov	62d9ed5ce7	Implement shortest-path routing for inference (#362 ) This PR: 1. Adds shortest path routing for inference. We build a graph with client-server and server-server latencies and compute costs, as well as empirically measured overheads. For client-server latencies, we ping possible first and last servers in a sequence in `SequenceManager.update()`. We penalize servers who may not have enough cache for our request. This uses info added to DHT in #355, #356, #358. 2. Makes a server ping neighboring servers in addition to next ones. This is to get an opportunity to change the server even before we use all its blocks (e.g., because a neighboring server is faster). This feature is not enabled though, since it increases graph size for N servers to O(N^2) - but we may enable it if needed. 3. Fixes a `SequenceManager` bug with the first `update()`. Previously, this update was likely to produce incorrect information and cause to `MissingBlocksErrors` until the next update happens.	10 months ago
Alexander Borzunov	11f0d992d7	Report inference, forward, and network RPS separately (#358 ) Inference RPS may be very different from forward RPS. E.g., currently bnb uses a completely different algorithm for NF4 inference. We report detailed RPS info that can be then used for shortest-path routing for inference.	10 months ago
Alexander Borzunov	81c4a45ca2	Make a server ping next servers (#356 ) This PR makes a server ping potential next servers in a chain and report the RTTs to DHT. This will be used for shortest-path routing.	10 months ago
Alexander Borzunov	2c8959e713	Share more info about a server in DHT (#355 )	10 months ago
justheuristic	37fdcb3fe0	Switch adapters slightly faster (#353 ) Currently, each `TransformerBackend.inference_step` looks for adapters and sets the correct adapter type for each block. This is not very expensive, but it can measurably affect inference time. This pull request uses faster adapter switching with just one variable assignment, without iterating over block.modules().	10 months ago
Alexander Borzunov	9703358df0	Fix bugs in _choose_num_blocks() added in #346 (#354 )	10 months ago
Alexander Borzunov	1a78638c02	Test that bitsandbytes is not imported when it's not used (#351 ) We avoid importing bitsandbytes when it's not used, since bitsandbytes doesn't always find correct CUDA libs and may raise exceptions because of that.	10 months ago
justheuristic	010857a834	Estimate adapter memory overhead in choose_num_blocks() (#346 ) * estimate adapter memory overhead * reduce number of heads based on that --------- Co-authored-by: Alexander Borzunov <borzunov.alexander@gmail.com>	10 months ago
Artem Chumachenko	b9f0a5467f	Support peft LoRA adapters (#335 ) Implement an option to deploy PEFT adapters to a server. Clients can set active_adapter=... to use these adapters. --------- Co-authored-by: Aleksandr Borzunov <borzunov.alexander@gmail.com> Co-authored-by: justheuristic <justheuristic@gmail.com>	10 months ago
Alexander Borzunov	fa095f6461	Use 4-bit for llama by default, use bitsandbytes 0.40.0.post3 (#340 ) NF4 inference with bitsandbytes 0.40.0.post3 is ~2x faster than int8 inference, though training is still ~3x slower, see: - [bitsandbytes 0.40.0 Release notes](https://github.com/TimDettmers/bitsandbytes/releases/tag/0.40.0) - [RPS benchmarks](https://github.com/bigscience-workshop/petals/pull/333#issuecomment-1614040385) We've decided to use NF4 by default for LLaMA.	10 months ago
Alexander Borzunov	158013a671	Implement direct server-to-server communication (#331 ) Implement #226.	10 months ago
Alexander Borzunov	de930918a0	Support loading blocks in 4-bit (QLoRA NF4 format, disabled by default) (#333 )	11 months ago
Alexander Borzunov	d126ee3053	Add benchmark scripts (#319 ) This PR: - Adds benchmark scripts for inference, forward pass, and full training step (e.g. used for experiments in our paper). - Fixes bug with dtypes in `petals.DistributedBloomForSequenceClassification`. - (minor refactor) Moves `DTYPE_MAP` to `petals.constants` as a useful constant.	11 months ago
Alexander Borzunov	cb3f018f9f	Add LLaMA support (#323 ) This PR: 1. Abolishes the model conversion procedure. Now, models are downloaded directly from original repositories like https://huggingface.co/bigscience/bloom. Servers download only shards with blocks to be hosted, and clients download only shards with input/output embeddings and layernorms. - BLOOM is loaded from `bigscience/bloom`, but we use the DHT prefix `bigscience/bloom-petals` for backward compatibility. Same with smaller BLOOMs and BLOOMZ. - LLaMA can be loaded from any repo like `username/llama-65b-hf`, but we use the DHT prefix `llama-65b-hf` (without the username) to accomodate blocks from different repos (there're a few of them with minor differences, such as `Llama` vs. `LLaMA` in the class name). 2. Refactors the client to generalize it for multiple models. Now, we have `petals.models` packages that contain model-specific code (e.g. `petals.models.bloom`, `petals.models.llama`). General code (e.g. CPU-efficient LM head, p-tuning) is kept in `petals.client`. 3. Introduces `WrappedLlamaBlock`, `DistributedLlamaConfig`, `DistributedLlamaForCausalLM`, `DistributedLlamaForSequenceClassification`, and `DistributedLlamaModel` compatible with Petals functionality (p-tuning, adapters, etc.). 4. Introduces `AutoDistributedConfig` that automatically chooses the correct config class (`DistributedLlamaConfig` or `DistributedBloomConfig`). The refactored configs contain all model-specific info for both clients and servers. Upgrade instructions: - Remove disk caches for blocks in old (converted) format to save disk space. That is, remove `~/.cache/petals/model--bigscience--bloom-petals` and `~/.cache/petals/model--bigscience--bloomz-petals` directories (if present).	11 months ago
Max Ryabinin	5c0733711a	Use number of tokens for attn_cache_size (#286 ) * Use number of tokens for attn_cache_size * Fix cache_bytes_per_block * Rename attn_cache_size to attn_cache_tokens	11 months ago
Max Ryabinin	c839173e57	Determine block dtype in a unified manner (#325 ) * Extract backend_dtype, remove duplicate DTYPE_MAP * Use bfloat16 as the default dtype, resolve dtype in load_pretrained_block	11 months ago
Max Ryabinin	3e7ae5116d	Remove unused imports and attributes (#324 ) * Remove unused imports and attributes	11 months ago
Alexander Borzunov	d9e7bfc949	Divide compute throughput by average no. of used blocks (#314 ) See #192.	1 year ago
Alexander Borzunov	21c3526ec1	Start SequenceManager's thread only after first .make_sequence() (#301 ) Why? - We'd like to avoid excess threads for the original sequence manager in case if we only use its slices (e.g. when we add adapters or need only a subset of model blocks): - If we create a sequence manager just before a fork (e.g. in a web app backend or a multi-thread benchmark), we'd like to avoid excess threads in the original process and only use this thread in child processes where we actually call `.make_sequence()`.	1 year ago
Alexander Borzunov	2116df08bc	Fix deps, enable 8-bit by default for TP (#298 ) This PR fixes issues of #290: - hivemind bfloat16 codec crashed on dummy tensors (with 0 elements), see https://github.com/learning-at-home/hivemind/pull/560 (this PR makes Petals depend on the latest hivemind version from the repo, it's temporary) - transformers version check mismatched with the version allowed in `setup.cfg` Also: - This PR enables 8-bit by default for TP. Even though TP in 8-bit may be slower, we currently prefer to host more blocks to increase the network's stability.	1 year ago
Alexander Borzunov	fee19e9b9b	Use get_logger(__name__) instead of get_logger(__file__) (#265 )	1 year ago
Alexander Borzunov	38b071135b	Show visible maddrs for public swarm too (#263 )	1 year ago
Alexander Borzunov	2a5070aa1a	Improve reachability logs (#253 )	1 year ago
justheuristic	c4938bc23e	Merge inference pools into one to increase inference speed (#225 ) It turns out using a separate pool for each block has led to significant slowdown, see #224 for details.	1 year ago
Alexander Borzunov	af3da5bb04	Choose --num_blocks automatically for all models (#217 )	1 year ago
Alexander Borzunov	6b12b0d050	Report server version and dht.client_mode in rpc_info(), check for updates on startup (#209 ) This PR: 1. Shows the current Petals version and checks for updates on startup. 2. Reports the current version and DHT mode in `rpc_info()`, so it can be shown on http://health.petals.ml or used on clients for efficient routing.	1 year ago
justheuristic	771ca590e7	Add service checking direct reachability from peers (#195 ) Servers joining from behind NATs/firewalls usually take several minutes to join a libp2p relay before they become accessible from the outside Internet. Moreover, requests to such servers are slower and more likely to fail (e.g., if the server switches a relay at the moment). If such servers host certain DHT keys, the swarm may occasionally lose read/write access to these keys, which results in: - Clients being unable to find any servers hosting a certain block. - All servers starting rebalancing to the same place to close the alleged "gap" in the swarm. This PRs modifies servers so that DHT keys are only hosted on directly reachable servers (the ones who aren't behind NAT/firewall). This way, DHT becomes more stable and works faster. Of course, trhe servers behind NATs/firewalls still accept requests for running inference/forward/backward for blocks they hold (it's more acceptable for this kind of requests to be slower or fail). Co-authored-by: Alexander Borzunov <borzunov.alexander@gmail.com>	1 year ago
Alexander Borzunov	a617ce3cfa	Fix psutil-related AccessDenied crash, disable --load_in_8bit by default in case of TP (#188 ) * Don't count open fds since it leads to AccessDenied crashes on some machines * Use --load_in_8bit=False by default in case of tensor parallelism * Install petals from PyPI in fine-tuning tutorials	1 year ago
Egiazarian Vage	93bed7da5a	Support libp2p relays for NAT traversal (#186 ) - Added relay options to servers - Enabled relay options by default - Changed hivemind version to 1.1.5 - Moved reachability check to be performed after blocks are loaded Co-authored-by: Alexander Borzunov <borzunov.alexander@gmail.com>	1 year ago
justheuristic	ae9e71fe8e	Add local tensor-parallel fwd/bwd (#143 ) This pull request adds an option to run Petals server on multiple local GPUs. It uses https://github.com/BlackSamorez/tensor_parallel - 8bit approximation error same as in main (mean~=2% q0.9~=5%) - TP=1, 2, 3 (see screenshots above) - forward, grad w.r.t. input and inference exact match with main with TP=1 - `>=`80% GPU utilization with 3x 1080ti, batch = 8 tokens - throughput measured with and without TP - TP on 1080Tis has near-linear speedup comparable to the benchmarks (see first message) Co-authored-by: Iaroslav Lisniak <yalisnyak@nes.ru> Co-authored-by: Andrei Panferov <andrei@blacksamorez.ru> Co-authored-by: Alexander Borzunov <borzunov.alexander@gmail.com>	1 year ago

1 2

66 Commits (5ce4f1a1598b1fca9fe6bd30cfbd85aa99bce2c7)