petals

Commit Graph

Author	SHA1	Message	Date
Alexander Borzunov	0a313bf6c5	Update hivemind to 1.1.8, enable efficient bfloat16 encoding (#311 ) This PR: 1. Updates hivemind to 1.1.8 (includes https://github.com/learning-at-home/hivemind/pull/565) 2. Enables efficient bfloat16 serialization by default (`USE_LEGACY_BFLOAT16 = False`) 3. Removes logging code that was included to hivemind in https://github.com/learning-at-home/hivemind/pull/542	1 year ago
Alexander Borzunov	8f6342a861	Refactor RemoteSequenceManager (#309 ) This PR: 1. Extracts `SequenceManagerConfig` and `SequenceManagerState` subclasses. The config is provided by caller and never changed from inside `RemoteSequenceManager`. The state is a part of the `RemoteSequenceManager`'s state shared between the main manager and its slices. We fix some slicing bugs along the way. 2. Removes `dht_prefix` and `p2p` arguments, makes `dht` argument optional. `dht_prefix` can always be overridden using `config.dht_prefix`. `p2p` actually needed only under the hood of `RemoteSequenceManager`, so it can extract it by itself without exposing this low-level class to callers. If strictly necessary, a caller can provide `p2p` as a part of `SequenceManagerState`. `dht` is also needed only by `RemoteSequenceManager`, so we can make it optional in the parent classes and create it automatically when it's not provided. 3. Simplifies retry logic. Previously, we could have "nested" retry loops: one in `._update()`, another in inference/forward/backward steps. The loop in `._update()` could introduce issues to concurrent inference/forward/backward calls, since it blocks the entire class if its delay period becomes too high. Now this logic is simplified: `._update()` performs only one attempt to fetch the DHT info, any retries are triggered by the inference/forward/backward steps. 4. Removes deprecated `RemoteTransformerBlock`. `RemoteTransformerBlock` was deprecated a long time ago, before Petals 1.0.0. Its removal is long due. 5. Removes `dht_utils.get_remote_module()`, `dht_utils.get_remote_sequence()`. This functions duplicate the functionality of the `RemoteSequential` constructor. 6. (minor) Removes `RemoteSequential.is_subsequence` flag. This flag worked incorrectly and was never used. I am removing it for the sake of simplicity.	1 year ago
Alexander Borzunov	454c193863	Fix OOMs happening in case of accelerate >= 0.16.0 (#310 ) - After #285, `load_pretrained_block()` uses `accelerate.utils.set_module_tensor_to_device()` - In accelerate>=0.16.0, it saves the tensor in the dtype previously used by the model instead of dtype of the weights (https://github.com/huggingface/accelerate/pull/920) - Because of that, blocks and attention caches used float32, which caused OOMs - This PR makes `load_pretrained_block()` respect `torch_dtype` (default: `"auto"`, which means reading `torch_dtype` from `config.json`)	1 year ago
Alexander Borzunov	93c4eba5d1	Bump version to 1.1.4 (#306 )	1 year ago
Alexander Borzunov	c0e0e1319d	Force transformers to use config.torch_dtype by default (#307 )	1 year ago
Alexander Borzunov	98be9ffe4c	Relax the rest of Hugging Face dependencies (#305 )	1 year ago
Alexander Borzunov	35662b4a16	Require bitsandbytes == 0.38.0.post2, hivemind == 1.1.7 (#302 ) In particular, this PR fixes 8-bit support on nvidia16 GPUs (such as 1660) by including https://github.com/TimDettmers/bitsandbytes/pull/292. This support was requested multiple times on Discord.	1 year ago
Alexander Borzunov	21c3526ec1	Start SequenceManager's thread only after first .make_sequence() (#301 ) Why? - We'd like to avoid excess threads for the original sequence manager in case if we only use its slices (e.g. when we add adapters or need only a subset of model blocks): - If we create a sequence manager just before a fork (e.g. in a web app backend or a multi-thread benchmark), we'd like to avoid excess threads in the original process and only use this thread in child processes where we actually call `.make_sequence()`.	1 year ago
Alexander Borzunov	6c6150f684	Remove use_auto_relay=True in client (#300 ) `use_auto_relay=True` makes the libp2p daemon look for relays to become reachable if we are behind NAT/firewall. However, being reachable is not necessary for the Petals client, and we should not spend the relays' capacity on this.	1 year ago
Alexander Borzunov	892fa2386a	Remove CustomLinear8bitLt (#297 ) This became a part of https://github.com/TimDettmers/bitsandbytes/releases/tag/0.37.0.	1 year ago
Alexander Borzunov	2116df08bc	Fix deps, enable 8-bit by default for TP (#298 ) This PR fixes issues of #290: - hivemind bfloat16 codec crashed on dummy tensors (with 0 elements), see https://github.com/learning-at-home/hivemind/pull/560 (this PR makes Petals depend on the latest hivemind version from the repo, it's temporary) - transformers version check mismatched with the version allowed in `setup.cfg` Also: - This PR enables 8-bit by default for TP. Even though TP in 8-bit may be slower, we currently prefer to host more blocks to increase the network's stability.	1 year ago
justheuristic	987f4d2b2f	Update bitsandbytes, hivemind, transformers (#290 ) - new bitsandbytes supports newer and older GPUs - new hivemind supports a better bfloat16 codec Co-authored-by: Alexander Borzunov <borzunov.alexander@gmail.com>	1 year ago
Alexander Borzunov	e0cef73757	Hotfix: Increase daemon_startup_timeout (#292 ) For some reasons, right now 15 sec is not enough to connect to the bootstrap peers in the public swarm, as reported by multiple users and observed by me. Increasing it to 120 sec until we find the root cause of the issue.	1 year ago
Max Ryabinin	793726b041	Speed up loading blocks using init with meta weights (#285 ) * Init WrappedBloomBlock with meta weights --------- Co-authored-by: Alexander Borzunov <borzunov.alexander@gmail.com>	1 year ago
Alexander Borzunov	c519bffc59	Bump version to 1.1.3 (#278 )	1 year ago
Alexander Borzunov	aae1f4f368	Increase default request_timeout (#276 ) This PR increases `request_timeout`, since the previous default of 30 sec is not enough for many use cases. Previously, we kept the request timeout low since we assumed that the server could freeze on dial if the target peer is behind a firewall. However, apparently, it won't freeze because libp2p has its own [dial timeout](https://github.com/libp2p/go-libp2p/blob/v0.26.0/core/network/context.go#L11).	1 year ago
justheuristic	fb2583b682	Use inference mode in _MergedInferenceStep (#275 )	1 year ago
Alexander Borzunov	fd9400b392	Fix use_chunked_forward="auto" on non-x86_64 machines (#267 ) Import of cpufeature may crash on non-x86_64 machines, so this PR makes the client import it only if necessary.	1 year ago
Alexander Borzunov	a2e7f27a5a	Improve "connect your GPU" message (#266 )	1 year ago
Alexander Borzunov	fee19e9b9b	Use get_logger(__name__) instead of get_logger(__file__) (#265 )	1 year ago
Alexander Borzunov	55e7dc07a0	Limit max delay between retries to 15 min (#264 )	1 year ago
Alexander Borzunov	38b071135b	Show visible maddrs for public swarm too (#263 )	1 year ago
Alexander Borzunov	2a5070aa1a	Improve reachability logs (#253 )	1 year ago
Alexander Borzunov	4091db10bf	Lower payload size threshold for stream handlers (#251 ) Hotfix: we add "// 2" since hivemind==1.1.5 serializes bfloat16 tensors in float32, so they take 2x more space.	1 year ago
Alexander Borzunov	9954cb84fe	Add `allowed_servers`, `max_retries` options to the client, improve logs (#235 )	1 year ago
Alexander Borzunov	3c523ab0d2	Fix TP crashing when hypo_ids are used (#249 )	1 year ago
Alexander Borzunov	b03efb1ef5	Bump version to 1.1.2 (#244 )	1 year ago
Artem Chumachenko	d4c687daca	Fix dtype error in fine-tuning notebooks (#231 )	1 year ago
justheuristic	c4938bc23e	Merge inference pools into one to increase inference speed (#225 ) It turns out using a separate pool for each block has led to significant slowdown, see #224 for details.	1 year ago
Shuchang Zhou	3189b395f0	Fix a typo in error message (#227 ) By the code context, it can be inferred that do_sample==False when control reaches this point.	1 year ago
Alexander Borzunov	af3da5bb04	Choose --num_blocks automatically for all models (#217 )	1 year ago
Alexander Borzunov	cea83d3356	Bump version to 1.1.1 (#214 )	1 year ago
Alexander Borzunov	825f5dbf2d	CI: Convert model only when convert_model.py or setup.cfg change (#213 ) This reduces the test running time by 2 times, unless convert_model.py or setup.cfg are changed.	1 year ago
Alexander Borzunov	5ff250bee9	Improve errors in case of missing blocks, suggest to join your own server (#212 )	1 year ago
Alexander Borzunov	6ba63c6cc8	Fix output shape when resuming generation (#211 ) Before this PR, `model.generate()` returned one excess token when resuming generation with an existing (the last token of the previous session, `session.last_token_id`). This is an unexpected behavior not convenient for the downstream apps, so this PR changes it until it's too late.	1 year ago
Alexander Borzunov	cc5e5d32c0	Don't switch blocks if it makes swarm disjoint (#210 ) Even if the swarm seems to have at least 2 servers for each block, turning off on one of the servers could break it. That's because once a server is turned off, others may move to a better position, creating a significant downtime on their way. This PR prohibits switching blocks if it would make the swarm disjoint along the way.	1 year ago
Alexander Borzunov	6b12b0d050	Report server version and dht.client_mode in rpc_info(), check for updates on startup (#209 ) This PR: 1. Shows the current Petals version and checks for updates on startup. 2. Reports the current version and DHT mode in `rpc_info()`, so it can be shown on http://health.petals.ml or used on clients for efficient routing.	1 year ago
justheuristic	771ca590e7	Add service checking direct reachability from peers (#195 ) Servers joining from behind NATs/firewalls usually take several minutes to join a libp2p relay before they become accessible from the outside Internet. Moreover, requests to such servers are slower and more likely to fail (e.g., if the server switches a relay at the moment). If such servers host certain DHT keys, the swarm may occasionally lose read/write access to these keys, which results in: - Clients being unable to find any servers hosting a certain block. - All servers starting rebalancing to the same place to close the alleged "gap" in the swarm. This PRs modifies servers so that DHT keys are only hosted on directly reachable servers (the ones who aren't behind NAT/firewall). This way, DHT becomes more stable and works faster. Of course, trhe servers behind NATs/firewalls still accept requests for running inference/forward/backward for blocks they hold (it's more acceptable for this kind of requests to be slower or fail). Co-authored-by: Alexander Borzunov <borzunov.alexander@gmail.com>	1 year ago
justheuristic	5f58f00649	Return available cache size in rpc_info() (#191 ) This PR makes servers return their free cache (in tokens * layers to make it compression-agnostic) To be used when calling make_sequence(optimize="inference")	1 year ago
justheuristic	012f840f7e	Use length-weighted sampling in routing for inference (#204 ) This pull-request implements a simple (1) greedy (2) latency-agnostic routing optimization that should speed up both our use cases. Why this exists: our effort to merge full routing (ping-aware, throughut-aware, dijkstra) is in a sorry state between several branches; merging it into main would take many days. Co-authored-by: Aleksandr Borzunov <borzunov.alexander@gmail.com>	1 year ago
Alexander Borzunov	42d1bbb568	Fix --no_auto_relay help (#199 )	1 year ago
Alexander Borzunov	b4f3224cda	Make client ignore blacklist if all servers holding a block are blacklisted (#197 ) If all servers holding a certain block are blacklisted, we should display errors from them instead of raising `No peers holding blocks`. Indeed, if the error is client-caused, the client should learn its reason from the latest error messages. In turn, if the error is server/network-caused and we only have a few servers, we'd better know the error instead of banning all the servers and making the user think that no servers are available.	1 year ago
Alexander Borzunov	127cf66bee	Ignore network RPS if we failed to measure it (#198 )	1 year ago
Alexander Borzunov	82c9f93ce6	Bump version to 1.1.0 (#190 )	1 year ago
Alexander Borzunov	a617ce3cfa	Fix psutil-related AccessDenied crash, disable --load_in_8bit by default in case of TP (#188 ) * Don't count open fds since it leads to AccessDenied crashes on some machines * Use --load_in_8bit=False by default in case of tensor parallelism * Install petals from PyPI in fine-tuning tutorials	1 year ago
Egiazarian Vage	93bed7da5a	Support libp2p relays for NAT traversal (#186 ) - Added relay options to servers - Enabled relay options by default - Changed hivemind version to 1.1.5 - Moved reachability check to be performed after blocks are loaded Co-authored-by: Alexander Borzunov <borzunov.alexander@gmail.com>	1 year ago
Alexander Borzunov	16b69d6050	Fix GiBs in the "insufficient disk space" message (#187 )	1 year ago
Alexander Borzunov	712f5a330f	Remove backup bootstrap peer	1 year ago
justheuristic	d1fa5eb260	hotfix: add initial peer that did not crash :) (#181 ) add hotfix initial peer (@borzunov's peers are down)	1 year ago
Alexander Borzunov	6dd9a938bd	Import bitsandbytes only if it's going to be used (#180 )	1 year ago
Alexander Borzunov	e27706358c	Use slightly less memory in .generate() (#177 )	1 year ago
Alexander Borzunov	55698381d0	Disable chunked_forward() on AVX512 CPUs (#179 )	1 year ago
Alexander Borzunov	6948a0c5ee	Allow to disable chunked forward (#176 )	1 year ago
justheuristic	ae9e71fe8e	Add local tensor-parallel fwd/bwd (#143 ) This pull request adds an option to run Petals server on multiple local GPUs. It uses https://github.com/BlackSamorez/tensor_parallel - 8bit approximation error same as in main (mean~=2% q0.9~=5%) - TP=1, 2, 3 (see screenshots above) - forward, grad w.r.t. input and inference exact match with main with TP=1 - `>=`80% GPU utilization with 3x 1080ti, batch = 8 tokens - throughput measured with and without TP - TP on 1080Tis has near-linear speedup comparable to the benchmarks (see first message) Co-authored-by: Iaroslav Lisniak <yalisnyak@nes.ru> Co-authored-by: Andrei Panferov <andrei@blacksamorez.ru> Co-authored-by: Alexander Borzunov <borzunov.alexander@gmail.com>	1 year ago
Aleksandr Borzunov	ff8ade8d3b	Bump version to 1.0.0	1 year ago
Alexander Borzunov	9997ada3bb	Shield alloc & free from cancellation (#163 ) A handler's RPC code may be cancelled due to a request timeout or a client closing the connection. Before this PR: - If `.cancel()` happens while waiting for `hivemind.utils.enter_asynchronously()`, the lock will never be released. - If `.cancel()` happens while doing that before freeing memory, the memory will never be freed. This PR fixes it by deferring the cancellation with [asyncio.shield()](https://docs.python.org/3/library/asyncio-task.html#asyncio.shield). Now, the cancellation will happen only when all locks are released and alloc/free has completed.	1 year ago
Alexander Borzunov	d6992fca63	Hot fix: Increase hivemind.P2P's startup_timeout for Colab, remove absent initial peer (#162 )	1 year ago
Alexander Borzunov	7cdc57a04b	Alloc inference cache as one contiguous buffer (#160 )	1 year ago
Alexander Borzunov	523a7cad33	Fix issues related to `petals` as a module (#159 ) 1. Added `from petals.client import *` to `petals/__init__.py`, so you can write just that: ```python from petals import DistributedBloomForCausalLM ``` I didn't do the same with server, since its classes are supposed to by used by `petals.cli.run_server`, not end-users. Though it's still possible to do `from petals.server.smth import smth` if necessary. 2. Fixed one more logging issue: log lines from hivemind were shown twice due to a bug in #156. 3. Removed unused `runtime.py`, since the server actually uses `hivemind.moe.Runtime`, and `runtime.py` has no significant changes comparing to it.	1 year ago
justheuristic	91898c3c90	Switch to speedtest-cli (#157 ) This pullrequest removes custom speed_test code in favour of speedtest-cli module. This is necessary to ensure that random warnings / print-outs do not mess with our outputs. Co-authored-by: Max Ryabinin <mryabinin0@gmail.com>	1 year ago
Alexander Borzunov	668b736031	Fix logging: do not duplicate lines, enable colors in Colab (#156 )	1 year ago
Alexander Borzunov	041ad20891	Check reachability automatically and give advice how to fix it (#155 ) 1. If we connect to the public swarm, the server now automatically checks its DHT's reachability from the outside world using API at http://health.petals.ml This is important to disallow unreachable servers to proceed (they create issues for the clients, such as repetitive retries). If http://health.petals.ml is down, the server proceeds without the check (so we don't depend on it). However, if health.petals.ml is up and explicitly tells us that we are unrechable, the server shows the reason of that and how to solve it. The check may be disabled with the `--skip_reachability_check` option (though I can't imagine cases where someone needs to use it). 2. Added `--port` and `--public_ip` as simplified convenience options for users not familiar with `--host_maddrs` and `--announce_maddrs`.	1 year ago
Alexander Borzunov	73df69a117	Reset MemoryCache during rebalancings (#154 ) Before this PR, if there were open inference sessions right when rebalancing is triggered, their cache was never properly destroyed.	1 year ago
Max Ryabinin	bd91be27ea	Add missing methods for SamplingAlgorithm, fix docstrings (#107 ) * Add missing methods for SamplingAlgorithm, fix docstrings * Add SamplingAlgorithm to _choose_sample_algorithm * Add test_sampling * Add a warning if sampling options were passed, but do_sample=False * Skip the sampling test for now Co-authored-by: Alexander Borzunov <borzunov.alexander@gmail.com>	1 year ago
Alexander Borzunov	701ec7e53e	Clean up disk space (#152 )	1 year ago
justheuristic	b04982c1a2	Bump transformers to 4.25.1 (#151 ) - latest accelerate, transformers, huggingface_hub - rearrange attention caches to support https://github.com/huggingface/transformers/pull/18344 - remove unused code - fix edge case where session crashes when receiving seq length 0 - assert transformer version when importing WrappedBloomBlock Co-authored-by: Alexander Borzunov <borzunov.alexander@gmail.com> Co-authored-by: Max Ryabinin <mryabinin0@gmail.com>	1 year ago
Alexander Borzunov	e4dc938dfe	Fix OOMs during server rebalancing (#150 ) The cause of OOMs were the cyclic references `TransformerBackend <-> PrioritizedTaskPool` that could not have been garbage collected properly. Still, I've added explicit tensor removal just in case.	1 year ago
Alexander Borzunov	83d9493b6c	Improve block size calculations (#149 )	1 year ago
Alexander Borzunov	84fec81543	Suppress asyncio error logs by default (#142 )	1 year ago
Alexander Borzunov	e99bf36647	Use common folder for all caches, make it a volume in Dockerfile (#141 )	1 year ago
Alexander Borzunov	e1d8793f00	Show route on client (#139 )	1 year ago
Alexander Borzunov	77a00e17f0	Fix "could not unlink the shared memory file" during rebalancing (#135 )	1 year ago
Alexander Borzunov	318d690a5c	Fix waiting until free memory is available (#136 )	1 year ago
Alexander Borzunov	e8fac92e59	Allow .generate() to reuse existing inference session (#132 )	1 year ago
Alexander Borzunov	1fe3716589	Don't ban servers in case of client-caused handler errors (#134 )	1 year ago
Alexander Borzunov	66f1799d32	Set default --step_timeout to 5 min (#133 )	1 year ago
Alexander Borzunov	f56edaa13f	Fix inference and rpc_info() fault tolerance (#131 )	1 year ago
justheuristic	79a4308992	Clear trigger before engaging in update (#130 ) Update sequence_manager.py	1 year ago
justheuristic	68c85e7492	Avoid synchronous updates, ban peers based on request outcome (#127 ) - sequence_manager now takes care for its own updated-ness - no need to manually update it - if a peer fails a request, sequence manager will ban this peer temporarily. Ban times increase with failure streaks Co-authored-by: Alexander Borzunov <borzunov.alexander@gmail.com>	1 year ago
Alexander Borzunov	9dbf5e2e6f	Set dht.num_workers = n_layer, update_period = 150, expiration = 300 (#125 )	1 year ago
Max Ryabinin	3ca8b4f082	Fix typos with codespell (#126 )	1 year ago
justheuristic	8491ed2bd3	Add checks for forward() inputs on the client side (#123 )	1 year ago
Max Ryabinin	055f85b83e	Call block.load_state_dict only once (#124 )	1 year ago
Alexander Borzunov	fc6722576b	Choose --num_blocks for bigscience/bloom-petals automatically (#119 )	1 year ago
Alexander Borzunov	f72c220404	Suppress quantization warning and fix dtype defaults in compute benchmark (#117 )	1 year ago
Alexander Borzunov	643a054170	Make server use smart defaults (#115 ) Summary: ```python parser.add_argument('--attn_cache_size', type=str, default=None, help='The size of GPU memory allocated for storing past attention keys/values between inference steps. ' 'Examples: 500MB, 1.2GB, 1073741824 (bytes). Note that 1KB != 1KiB here. ' 'Default: 0.5GiB * num_blocks * hidden_size / 14336. ' 'The latter is the hidden size of the bigscience/bloom-petals model.') parser.add_argument('--request_timeout', type=float, required=False, default=3 * 60, help='Timeout (in seconds) for the whole rpc_forward/rpc_backward/rpc_forward_stream/rpc_backward_stream request') parser.add_argument('--session_timeout', type=float, required=False, default=30 * 60, help='Timeout (in seconds) for the whole inference session') parser.add_argument('--step_timeout', type=float, required=False, default=60, help="Timeout (in seconds) for waiting the next step's inputs inside an inference session") parser.add_argument('--load_in_8bit', type=bool, default=None, help="Convert the loaded model into mixed-8bit quantized model. Default: True if GPU is available") ``` Co-authored-by: justheuristic <justheuristic@gmail.com>	1 year ago
justheuristic	9e11f73242	Fix tile size on ampere (#116 ) Fix tile size on ampere Co-authored-by: Aleksandr Borzunov <borzunov.alexander@gmail.com>	1 year ago
justheuristic	617d70f7dc	Support --load_in_8bit on pre-Turing GPUs (#113 ) - Linear8bitLt now supports for pre-turing GPUs by temporarily upcasting quantized weights. - added a test for linear8bitlt accuracy with the new fallback, the accuracy is similar than the real thing, (slightly better due to non-quantized A) - performance is roughly halfway between the default mode and memory_efficient_backward Alternatives considered: - cupy - slow, casting to float internally - triton - fast but unstable af. every 3rd attempt to matmul is a segfault - bnb.functional.igemm (no lt) - "CuBLAS Error 8" on old GPUs Co-authored-by: Aleksandr Borzunov <borzunov.alexander@gmail.com>	1 year ago
Alexander Borzunov	1ea44b0d3c	Measure throughput for different configs, devices, and dtypes separately (#114 )	1 year ago
justheuristic	01838f9a99	Fix Linear8bitlt state config, update tests (#112 ) * fix state initializer * update tests to actually use new code * keep bias during quantization	1 year ago
Aleksandr Borzunov	96033de921	Fix script for running servers robustly	1 year ago
Aleksandr Borzunov	85cf32d2a4	Add script to run servers robustly	1 year ago
justheuristic	088713912d	Patch Linear8bit to enable CxB backward (#111 ) A patch to bitsandbytes 0.34.0 that introduces an option to run backward pass in default (fast) matrix layout. Authors: cxb inversion by @borzunov, original 8bit code by @timdettmers * optimized layout inversion code by @borzunov ([original code](https://colab.research.google.com/drive/1EJ0MKifajXSSVq7O2_QGwtb0l6gRAGrh?usp=sharing)) to use less forward calls * implemented CustomLinear8bitLt, a child of Linear8bitLt that can do backward without CB * added exact match tests for layouts and linear layers: see tests/test_linear8bitlt.py * switched petals to the new layer type Core idea: layouts apply the same permutation to every tile in the matrix. We can treat this as (batched) gather ops. Reshape input tensor so that ij-th gather operation op will apply to ij-th elements in each tile. Prototype: Layout info: https://github.com/TimDettmers/bitsandbytes/blob/main/csrc/kernels.cu#L2130-L2136 Co-authored-by: Alexander Borzunov <hxrussia@gmail.com> Co-authored-by: Aleksandr Borzunov <borzunov.alexander@gmail.com> Co-authored-by: Tim Dettmers <tim.dettmers@gmail.com>	1 year ago
justheuristic	8dc0f513ba	Hotfix span selection (#110 ) Fix an issue in span selection that was introduced in #106	1 year ago
justheuristic	a2066a4096	Optimize RemoteSequenceManager (#106 ) - [x] made RemoteSequenceManager into a background thread that pre-fetches information instead of running just in time - [x] moved routing-related stuff to petals.client.routing - [x] extract remote peer routing information to RemoteSequenceInfo - [x] made sure that the code survives continued use (e.g. one hour) - [x] updated every spot where update_ is called manually - [x] modified get_sequence to check that the thread is alive, warn if not - [x] removed max_retries, switched rpc_info to exponential backoff - [x] fixed a bg that causes RemoteSeq* to lose user-defined hyperparameters (e.g. timeout) upon subsequencing (sequential[3:5]) - [x] moved client-side points strategy to client.routing - [x] ensured that RemoteSequenceManager thread created in get_remote_module properly shuts down when the module is destroyed - [x] resolved minor affected todos - [x] modified tests to no longer use PYTHONPATH - [x] worked around protocol error in rpc_info Co-authored-by: Aleksandr Borzunov <borzunov.alexander@gmail.com> Co-authored-by: Artem Chumachenko <artek.chumak@gmail.com>	1 year ago
Artem Chumachenko	7d859a947b	Expose request_timeout to DistributedBloomConfig (#105 ) Co-authored-by: Alexander Borzunov <borzunov.alexander@gmail.com>	1 year ago
Max Ryabinin	9faf08b898	Remove unused imports, add missing arguments to docstrings (#108 ) * Remove unused imports, add missing arguments to docstrings	1 year ago
justheuristic	b3115dac58	Update throughput.py	1 year ago
Alexander Borzunov	0a1cd3b9ba	Fix ptune with `low_cpu_mem_usage=True` (as in Colab) (#103 ) Fixes: - An exception while creating a model with `ptune/deep_ptune` and `low_cpu_mem_usage=True` (which is currently default). - dtype mismatch between the prompts and the rest of the model in `.forward()`.	1 year ago
Alexander Borzunov	43ac6016ac	Fix dtypes in backend schemas (#99 ) Currently, the schemas use `torch.float32`, so all inputs and outputs converted to float32 before sending and after receiving on both servers and clients. This creates a huge slowdown for the system. * This PR makes the schemas use the server's `--torch_dtype` argument (default is `torch.bloat16` for BLOOM-176B) * an option for client to request a specific output compression. Use case 1: client sends quantized inputs and expects quantized inputs in return. Use case 2: client uses quantization for gradients w.r.t. activations, but keeps grads w.r.t. __prompts__ as is for greater precision. * a comment explaining the purpose of NoSpendingPolicy - since we likely won't have it for the workshop * a test with custom compression (janky implementation for testing purposes) Co-authored-by: justheuristic <justheuristic@gmail.com>	1 year ago
Alexander Borzunov	7bd5916744	Make Petals a pip-installable package (attempt 2) (#102 ) 1. Petals can be now installed using `pip install git+https://github.com/bigscience-workshop/petals` - In case if you already cloned the repo, you can do `pip install .` or `pip install .[dev]` 2. Moved `src` => `src/petals` - Replaced `from src.smth import smth` with `from petals.smth import smth` 3. Moved `cli` => `src/petals/cli` - Replaced `python -m cli.run_smth` with `python -m petals.cli.run_smth` (all utilities are now available right after pip installation) 4. Moved the `requirements*.txt` contents to `setup.cfg` (`requirements.txt` for packages is not supported well by modern packaging utils) 5. Increased the package version from `0.2` to `1.0alpha1`	1 year ago

1 2 3 4 5

201 Commits (main)