petals

Commit Graph

Author	SHA1	Message	Date
Alexander Borzunov	b945e388e5	Remove smaller limit for legacy bfloat16 serialization	9 months ago
Max Ryabinin	1ebd88ae7b	Optimize the Falcon block for inference (#500 ) This PR attempts to optimize the inference of Falcon models in the single-token setup by reducing the majority of Python overhead and making several assumptions about the setup. Specifically, * Layer normalization, QKV projection (with splitting) and rotary embeddings are executed through CUDA graphs, which reduces most overhead related to small kernel launche * If no sin/cos tensors are cached by the rotary embedding layer, we cache them for 8192 tokens (INFERENCE_MAX_LENGTH) during the first forward pass. In general, it should be beneficial to always run a max-length sequence before starting a block, but this is a question for another PR The PR also adds a small test to ensure that the results (without quantization) of the block before and after quantization indeed match. Lastly, the pull request makes the backward pass work (as discussed in https://github.com/bigscience-workshop/petals/pull/499) by making cached sin/cos for RotaryEmbedding into buffers and disabling the inference mode during their creation.	9 months ago
Alexander Borzunov	d40eb6c701	Fix prompt tuning after #464 (#501 ) Unfortunately, running inference in models with `"ptune" in config.tuning_mode` was broken after #464.	9 months ago
Alexander Borzunov	dd4a3230bc	Add Falcon support (#499 ) This PR adds: - Support for models based on `transformers.FalconModel` (the in-library format for Falcon). Tested on Falcon-40B. - CI tests for Falcon-RW-1B. - `--throughput dry_run` option to evaluate throughput and exit right away (implemented by @mryab). Limitations: - Backward pass support is broken for now, will be fixed in #500. Co-authored-by: Max Ryabinin <mryabinin0@gmail.com>	9 months ago
Alexander Borzunov	b4d822afb2	Force use_cache=True in config only (#497 ) This reverts a part of #496 and instead overrides `use_cache` in `LlamaConfig`s only (so the correct value is visible by HF `.generate()` as well).	9 months ago
Alexander Borzunov	abd547735f	Force use_cache=True (#496 )	9 months ago
Alexander Borzunov	6ef6bf5fa2	Create model index in DHT (#491 ) This PR creates an index of models hosted in the swarm - it is useful to know which custom models users run and display them at https://health.petals.dev as "not officially supported" models.	9 months ago
Alexander Borzunov	6bb3f54e39	Replace dots in repo names when building DHT prefixes (#489 )	9 months ago
Alexander Borzunov	02fc71eb25	Fix race condition in MemoryCache (#487 )	9 months ago
Alexander Borzunov	dc0072fde1	Wait for DHT storing state OFFLINE on shutdown (#486 )	9 months ago
Alexander Borzunov	a26559ff65	Fix `.generate(input_ids=...)` (#485 )	9 months ago
Alexander Borzunov	459933f846	Remove no-op process in PrioritizedTaskPool (#484 ) Please revert this if you ever need to make `PrioritizedTaskPool` a process again.	9 months ago
Alexander Borzunov	26ebbfe8f0	Support macOS (#477 ) This PR makes both clients and servers work on macOS. Specifically, it: - Follows https://github.com/learning-at-home/hivemind/pull/586 to run a macOS-compatible `p2pd` binary (both x86-64 and ARM64 are supported) - Fixes forking issues and tests on macOS, Python 3.10+ - Introduces basic support for serving model blocks on Apple M1/M2 GPUs (torch.mps) - Increases max number of open files by default (it's not enough on Linux and is really small on macOS)	9 months ago
justheuristic	c08d09c4d3	Rewrite MemoryCache alloc_timeout logic (#434 ) - rpc_inference: server will now accept allocation timeout from user, defaults to no timeout - bugfix: inference timeout is now measured from the moment the request is received - previously, you would have to wait for your timeout plus the time it takes to sort through the queue (other users' timeout) - now, you get AllocationFailed if you had to wait for over (timeout) seconds - regardless of other users - a request for inference with no timeout will now fail instantly if there is not enough memory available - dtype number of bytes is now correctly determined for int, bool & other types --------- Co-authored-by: Your Name <you@example.com> Co-authored-by: Alexander Borzunov <borzunov.alexander@gmail.com> Co-authored-by: Aleksandr Borzunov <hxrussia@gmail.com>	9 months ago
Alexander Borzunov	90840dfea2	Fix requiring transformers>=4.32.0 (#480 )	9 months ago
Alexander Borzunov	915b357740	Require transformers>=4.32.0 (#479 ) It's necessary to load https://huggingface.co/petals-team/StableBeluga2 since it doesn't have deprecated `inv_freq` weights.	9 months ago
Alexander Borzunov	6967904590	Bump version to 2.1.0 (#474 ) * Bump version to 2.1.0 * Suggest using resharded repo * LLaMA -> Llama in readme	9 months ago
Alexander Borzunov	df8ab09ca2	Hide excess key message (#476 ) Before: ```python Aug 23 23:51:31.394 [INFO] Loaded Maykeye/TinyLLama-v0 block 0, _IncompatibleKeys(missing_keys=[], unexpected_keys=['self_attn.rotary_emb.inv_freq']) ``` After: ```python Aug 23 23:51:31.394 [INFO] Loaded Maykeye/TinyLLama-v0 block 0 ``` Hiding this since the excess keys in Llama-based models are okay since the latest transformers release.	9 months ago
Artem Chumachenko	a14ae7334d	Update peft to 0.5.0 version (#475 ) Update peft to 0.5.0	9 months ago
Alexander Borzunov	a9b0e9ff1a	Support loading weights from Safetensors on server (#473 )	9 months ago
justheuristic	9250025140	Support transformers 4.32.x (#471 )	9 months ago
Alexander Borzunov	de2475f31c	Make client compatible with transformers' GenerationMixin (#464 ) This PR drops custom generation codes and introduces compatibility with `transformers.GenerationMixin` instead. This includes support for more sampling options (`top_p`, `top_k`, `repetition_penalty` requested in #460) and beam search - all that is now identical to running model with transformers locally. Most features (excluding beam search and other rarely used stuff) are also compatible with resuming existing sessions. ### Breaking changes If `.generate()` or forward passes are being run inside an `.inference_session()` context, they now use the opened session by default. So, these snippets are now equivalent: ```python # Using default session with model.inference_session(max_length=100): output_ids = model.generate(input_ids, max_new_tokens=3) # Explicitly specifying a session with model.inference_session(max_length=100) as sess: output_ids = model.generate(input_ids, max_new_tokens=3, session=sess) ``` Earlier, the 1st snippet was creating a new session, which is not what most people expected (= such code was most likely to introduce a bug, which is now fixed).	10 months ago
Alexander Borzunov	063e94b4c8	Move SequenceManagerConfig -> ClientConfig, petals.dht_utils -> petals.utils.dht (#463 )	10 months ago
Artem Chumachenko	568f21dc3b	Add customizable input tensors (#445 )	10 months ago
Alexander Borzunov	329f7d31e8	Add `blocked_servers` argument (#462 ) Should be used as: ```python model = AutoDistributedModelForCausalLM(model_name, blocked_servers=[peer_id1, peer_id2]) ```	10 months ago
Alexander Borzunov	722c4dc496	Bump version to 2.0.1.post2 (#459 )	10 months ago
Alexander Borzunov	056f22515a	Prioritize short inference, unmerge pools for long inference (#458 ) Right now, long inference requests may occupy Runtime for a few seconds without giving it away to process short (most latency-sensitive requests). This PR fixes it by disallowing the merged pool for long requests and prioritizing the short ones.	10 months ago
justheuristic	55eb36ef48	Fix missing torch.cuda.synchronize for computing throughput (#456 )	10 months ago
Alexander Borzunov	8c546d988a	Test Llama, rebalancing, throughput eval, and all CLI scripts (#452 ) This PR extends CI to: 1. Test Llama code using [TinyLlama-v0](https://huggingface.co/Maykeye/TinyLLama-v0). 2. Test rebalancing (sets up a situation where the 1st server needs to change its original position). 3. Check if benchmark scripts run (in case someone breaks its code). Note that the benchmark results are meaningless here (since they're measured on a tiny swarm of CPU servers, with low `--n_steps`). 4. Test `petals.cli.run_dht`. 5. Increase swap space and watch free RAM (a common issue is that actions are cancelled without explanation if there's not enough RAM - so it's a useful reminder + debug tool). 6. Fix flapping tests for bloom-560m by increasing tolerance. Other minor changes: fix `--help` messages to show defaults, fix docs, tune rebalancing constants.	10 months ago
Alexander Borzunov	df6fdd2d0b	Force using --new_swarm instead of empty --initial_peers (#451 ) This prohibits passing `--initial_peers` without arguments, since it's likely to be a side-effect from `--initial_peers $INITIAL_PEERS` with the env var not set. Users should use `--new_swarm` for that, as explained in the private swarm tutorial.	10 months ago
Alexander Borzunov	2a150770a4	Prefer longer servers for fine-tuning, exclude unreachable (#448 ) We choose longer servers to minimize the number of hops but leave some randomization to distribute the load. We also exclude servers known to be unreachable.	10 months ago
Alexander Borzunov	00d48dcbe1	Override float32 in config to bfloat16 (#431 )	10 months ago
justheuristic	ac9b546706	[Refactor] extract block forward, backward and inference into a separate file (#435 ) This PR does not change any functionality. It merely moves stuff around. List of changes: handler.py/_rpc_forward became block_methods/rpc_forward handler.py/_rpc_backward became block_methods/rpc_backward the math bits of rpc_inference were extracted into block_methods/iterate_rpc_inference --------- Co-authored-by: Your Name <you@example.com> Co-authored-by: artek0chumak <artek.chumak@gmail.com> Co-authored-by: Aleksandr Borzunov <borzunov.alexander@gmail.com>	10 months ago
Vadim Peretokin	d0b5af34cd	Fix typo and make blocks message more informative (#437 ) The message really doesn't tell me much as a user, since I never touched update_period to begin with: ``` Aug 06 09:43:07.287 [WARN] [petals.server.server.run:701] Declaring blocs to DHT takes more than --update_period, consider increasing it ``` Made it better and more informative.	10 months ago
Alexander Borzunov	a1f7791d5e	Fix petals.utils.ping for servers with client-mode DHT (#430 ) Fix #429.	10 months ago
Alexander Borzunov	351e96bc46	Penalize servers that use relays during rebalancing (#428 ) Servers accessible only via relays may introduce issues if they are the only type of servers holding certain blocks. Specifically, a connection to such servers may be unstable or opened after a certain delay. This PR changes their self-reported throughput, so that the rebalancing algorithm prefers to put directly available servers for hosting each block.	10 months ago
Alexander Borzunov	44fefa5e54	Add connect_timeout (#423 )	10 months ago
Alexander Borzunov	f3fafd14a4	Bump version to 2.0.1 (#411 )	10 months ago
Alexander Borzunov	fd19c21859	Update --update_period and --expiration defaults (#410 )	10 months ago
justheuristic	5af04524dd	Split long sequences into chunks (#403 ) This PR is designed to avoid OOMs when processing long sequences that happen due to the huge attention logits matrices. Co-authored-by: Alexander Borzunov <borzunov.alexander@gmail.com>	10 months ago
Alexander Borzunov	30b94ef18b	If speedtest fails, assume network speed of 100 Mbit/s (#404 ) The value is chosen as some safe value below average at https://health.petals.dev/ Note that if a server uses relays, the effective throughput will be further divided by 2 (see #399).	10 months ago
Alexander Borzunov	8666653cf5	Fix routing through relay, default network RPS, --token, logging, readme (#399 ) * Hide GeneratorExit in _iterate_inference_steps() * Update README.md about `--public_name` * Use .from_pretrained(..., use_auth_token=token) instead of token=token until it's fully supported across HF libs * Use default network speed 25 Mbit/s * Apply relay penalty in max-throughput routing * Replace RPS with "tokens/sec per block" in logs * Increase default expiration	10 months ago
Alexander Borzunov	6e4ebb94d2	Fix deadlocks in MemoryCache (#396 ) - Fix deadlocks in MemoryCache - Set default --alloc_timeout to 1 until the MemoryCache update	11 months ago
Alexander Borzunov	b6b3ae964f	Fix --attn_cache_tokens default (#392 )	11 months ago
Alexander Borzunov	d49d9ad0cf	Bump version to 2.0.0.post3 (#391 )	11 months ago
justheuristic	e51e84631d	Update to petals.dev (#390 ) Since `petals.ml` DNS record is still unavailable, we're switching everything to https://petals.dev Co-authored-by: Aleksandr Borzunov <hxrussia@gmail.com>	11 months ago
Aleksandr Borzunov	ddcda02b06	Hardcode IPs until DNS issues get resolved	11 months ago
Alexander Borzunov	b1ff8bdd6c	Bump version to 2.0.0.post1 (#384 )	11 months ago
Alexander Borzunov	057a2fb5de	Support Llama 2 (#379 )	11 months ago
Alexander Borzunov	3218534745	Fix --token arg (#378 )	11 months ago

1 2 3 4

188 Commits (payload-size)