petals

Commit Graph

Author	SHA1	Message	Date
Alexander Borzunov	2116df08bc	Fix deps, enable 8-bit by default for TP (#298 ) This PR fixes issues of #290: - hivemind bfloat16 codec crashed on dummy tensors (with 0 elements), see https://github.com/learning-at-home/hivemind/pull/560 (this PR makes Petals depend on the latest hivemind version from the repo, it's temporary) - transformers version check mismatched with the version allowed in `setup.cfg` Also: - This PR enables 8-bit by default for TP. Even though TP in 8-bit may be slower, we currently prefer to host more blocks to increase the network's stability.	2 years ago
justheuristic	987f4d2b2f	Update bitsandbytes, hivemind, transformers (#290 ) - new bitsandbytes supports newer and older GPUs - new hivemind supports a better bfloat16 codec Co-authored-by: Alexander Borzunov <borzunov.alexander@gmail.com>	2 years ago
Alexander Borzunov	e0cef73757	Hotfix: Increase daemon_startup_timeout (#292 ) For some reasons, right now 15 sec is not enough to connect to the bootstrap peers in the public swarm, as reported by multiple users and observed by me. Increasing it to 120 sec until we find the root cause of the issue.	2 years ago
Alexander Borzunov	a7d3d02194	Fix invalid author email in setup.cfg (#287 )	2 years ago
Alexander Borzunov	8dab37c1a9	Add benchmarks to readme (#284 )	2 years ago
Max Ryabinin	793726b041	Speed up loading blocks using init with meta weights (#285 ) * Init WrappedBloomBlock with meta weights --------- Co-authored-by: Alexander Borzunov <borzunov.alexander@gmail.com>	2 years ago
Alexander Borzunov	c519bffc59	Bump version to 1.1.3 (#278 )	2 years ago
Alexander Borzunov	aae1f4f368	Increase default request_timeout (#276 ) This PR increases `request_timeout`, since the previous default of 30 sec is not enough for many use cases. Previously, we kept the request timeout low since we assumed that the server could freeze on dial if the target peer is behind a firewall. However, apparently, it won't freeze because libp2p has its own [dial timeout](https://github.com/libp2p/go-libp2p/blob/v0.26.0/core/network/context.go#L11).	2 years ago
justheuristic	fb2583b682	Use inference mode in _MergedInferenceStep (#275 )	2 years ago
Alexander Borzunov	fd9400b392	Fix use_chunked_forward="auto" on non-x86_64 machines (#267 ) Import of cpufeature may crash on non-x86_64 machines, so this PR makes the client import it only if necessary.	2 years ago
Alexander Borzunov	a2e7f27a5a	Improve "connect your GPU" message (#266 )	2 years ago
Alexander Borzunov	fee19e9b9b	Use get_logger(__name__) instead of get_logger(__file__) (#265 )	2 years ago
Alexander Borzunov	55e7dc07a0	Limit max delay between retries to 15 min (#264 )	2 years ago
Alexander Borzunov	38b071135b	Show visible maddrs for public swarm too (#263 )	2 years ago
Alexander Borzunov	42594e5173	Link FAQ in readme (#260 )	2 years ago
Alexander Borzunov	2a5070aa1a	Improve reachability logs (#253 )	2 years ago
Alexander Borzunov	4091db10bf	Lower payload size threshold for stream handlers (#251 ) Hotfix: we add "// 2" since hivemind==1.1.5 serializes bfloat16 tensors in float32, so they take 2x more space.	2 years ago
Alexander Borzunov	9954cb84fe	Add `allowed_servers`, `max_retries` options to the client, improve logs (#235 )	2 years ago
Alexander Borzunov	3c523ab0d2	Fix TP crashing when hypo_ids are used (#249 )	2 years ago
justheuristic	b8a6788490	Fix examples/sst, add cls_model embeddings (#248 )	2 years ago
justheuristic	8766a14d28	Minor changes to examples/prompt-tuning notebooks (#247 ) Minor code changes required to run the notebook in a clean python environment	2 years ago
Alexander Borzunov	5367523df8	Fix typo in prompt-tuning-sst2.ipynb (#245 )	2 years ago
Alexander Borzunov	b03efb1ef5	Bump version to 1.1.2 (#244 )	2 years ago
Alexander Borzunov	5d7395e1b5	Prompt-tuning notebooks: suggest to use a smaller model for faster prototyping (#234 )	2 years ago
Artem Chumachenko	d4c687daca	Fix dtype error in fine-tuning notebooks (#231 )	2 years ago
Muhtasham Oblokulov	0ebf6de117	Add citation to readme (#219 ) Co-authored-by: Alexander Borzunov <borzunov.alexander@gmail.com>	2 years ago
justheuristic	c4938bc23e	Merge inference pools into one to increase inference speed (#225 ) It turns out using a separate pool for each block has led to significant slowdown, see #224 for details.	2 years ago
Shuchang Zhou	3189b395f0	Fix a typo in error message (#227 ) By the code context, it can be inferred that do_sample==False when control reaches this point.	2 years ago
Alexander Borzunov	fa5ac6e3b4	Mention BLOOMZ in readme (#221 )	2 years ago
Alexander Borzunov	e651d73f11	Add one more link to the "Getting started" tutorial (#218 ) Some people miss the "Try now in Colab" link or don't understand that it leads to the comprehensive tutorial, so I added one more explicit link.	2 years ago
Alexander Borzunov	af3da5bb04	Choose --num_blocks automatically for all models (#217 )	2 years ago
Alexander Borzunov	cea83d3356	Bump version to 1.1.1 (#214 )	2 years ago
Alexander Borzunov	702bb5a2c2	CI: Update deprecated actions, don't measure network RPS (#215 ) * CI: Switch to actions/cache@v3 (v2 is deprecated) * Don't run measure_network_rps() in tests since it doesn't work well in CI	2 years ago
Alexander Borzunov	825f5dbf2d	CI: Convert model only when convert_model.py or setup.cfg change (#213 ) This reduces the test running time by 2 times, unless convert_model.py or setup.cfg are changed.	2 years ago
Alexander Borzunov	5ff250bee9	Improve errors in case of missing blocks, suggest to join your own server (#212 )	2 years ago
Alexander Borzunov	6ba63c6cc8	Fix output shape when resuming generation (#211 ) Before this PR, `model.generate()` returned one excess token when resuming generation with an existing (the last token of the previous session, `session.last_token_id`). This is an unexpected behavior not convenient for the downstream apps, so this PR changes it until it's too late.	2 years ago
Alexander Borzunov	cc5e5d32c0	Don't switch blocks if it makes swarm disjoint (#210 ) Even if the swarm seems to have at least 2 servers for each block, turning off on one of the servers could break it. That's because once a server is turned off, others may move to a better position, creating a significant downtime on their way. This PR prohibits switching blocks if it would make the swarm disjoint along the way.	2 years ago
Alexander Borzunov	6b12b0d050	Report server version and dht.client_mode in rpc_info(), check for updates on startup (#209 ) This PR: 1. Shows the current Petals version and checks for updates on startup. 2. Reports the current version and DHT mode in `rpc_info()`, so it can be shown on http://health.petals.ml or used on clients for efficient routing.	2 years ago
justheuristic	771ca590e7	Add service checking direct reachability from peers (#195 ) Servers joining from behind NATs/firewalls usually take several minutes to join a libp2p relay before they become accessible from the outside Internet. Moreover, requests to such servers are slower and more likely to fail (e.g., if the server switches a relay at the moment). If such servers host certain DHT keys, the swarm may occasionally lose read/write access to these keys, which results in: - Clients being unable to find any servers hosting a certain block. - All servers starting rebalancing to the same place to close the alleged "gap" in the swarm. This PRs modifies servers so that DHT keys are only hosted on directly reachable servers (the ones who aren't behind NAT/firewall). This way, DHT becomes more stable and works faster. Of course, trhe servers behind NATs/firewalls still accept requests for running inference/forward/backward for blocks they hold (it's more acceptable for this kind of requests to be slower or fail). Co-authored-by: Alexander Borzunov <borzunov.alexander@gmail.com>	2 years ago
justheuristic	5f58f00649	Return available cache size in rpc_info() (#191 ) This PR makes servers return their free cache (in tokens * layers to make it compression-agnostic) To be used when calling make_sequence(optimize="inference")	2 years ago
Alexander Borzunov	37373a66c3	Update Anaconda installation commands (#205 )	2 years ago
justheuristic	012f840f7e	Use length-weighted sampling in routing for inference (#204 ) This pull-request implements a simple (1) greedy (2) latency-agnostic routing optimization that should speed up both our use cases. Why this exists: our effort to merge full routing (ping-aware, throughut-aware, dijkstra) is in a sorry state between several branches; merging it into main would take many days. Co-authored-by: Aleksandr Borzunov <borzunov.alexander@gmail.com>	2 years ago
Alexander Borzunov	42d1bbb568	Fix --no_auto_relay help (#199 )	2 years ago
justheuristic	c2cb6d19ae	Increase tolerances in test_tp_block (#196 ) deflapify tests	2 years ago
Alexander Borzunov	b4f3224cda	Make client ignore blacklist if all servers holding a block are blacklisted (#197 ) If all servers holding a certain block are blacklisted, we should display errors from them instead of raising `No peers holding blocks`. Indeed, if the error is client-caused, the client should learn its reason from the latest error messages. In turn, if the error is server/network-caused and we only have a few servers, we'd better know the error instead of banning all the servers and making the user think that no servers are available.	2 years ago
Alexander Borzunov	127cf66bee	Ignore network RPS if we failed to measure it (#198 )	2 years ago
Alexander Borzunov	487411e87e	Fix fine-tuning notebooks intros (#194 ) The notebook intros were outdated and mentioned the 6B model, while the actual code already runs the 176B model. This led to confusion among our users in Discord.	2 years ago
Alexander Borzunov	82c9f93ce6	Bump version to 1.1.0 (#190 )	2 years ago
Alexander Borzunov	a617ce3cfa	Fix psutil-related AccessDenied crash, disable --load_in_8bit by default in case of TP (#188 ) * Don't count open fds since it leads to AccessDenied crashes on some machines * Use --load_in_8bit=False by default in case of tensor parallelism * Install petals from PyPI in fine-tuning tutorials	2 years ago
Egiazarian Vage	93bed7da5a	Support libp2p relays for NAT traversal (#186 ) - Added relay options to servers - Enabled relay options by default - Changed hivemind version to 1.1.5 - Moved reachability check to be performed after blocks are loaded Co-authored-by: Alexander Borzunov <borzunov.alexander@gmail.com>	2 years ago

... 2 3 4 5 6 ...

513 Commits (19be29e89e2fe15a68e225d8c44986b66a058b7e) All Branches Search

513 Commits (19be29e89e2fe15a68e225d8c44986b66a058b7e)

All Branches