Commit Graph

491 Commits

Author SHA1 Message Date
Alexander Borzunov
d6992fca63
Hot fix: Increase hivemind.P2P's startup_timeout for Colab, remove absent initial peer (#162) 2022-12-19 19:36:20 +04:00
Artem Chumachenko
0a6b5f31aa
Fix misstypos in the example notebooks. (#161)
Fix misstypos
2022-12-16 19:42:18 +04:00
Alexander Borzunov
7cdc57a04b
Alloc inference cache as one contiguous buffer (#160) 2022-12-16 10:38:20 +04:00
Alexander Borzunov
523a7cad33
Fix issues related to petals as a module (#159)
1. Added `from petals.client import *` to `petals/__init__.py`, so you can write just that:

    ```python
    from petals import DistributedBloomForCausalLM
    ```

    I didn't do the same with server, since its classes are supposed to by used by `petals.cli.run_server`, not end-users. Though it's still possible to do `from petals.server.smth import smth` if necessary.

2. Fixed one more logging issue: log lines from hivemind were shown twice due to a bug in #156.

3. Removed unused `runtime.py`, since the server actually uses `hivemind.moe.Runtime`, and `runtime.py` has no significant changes comparing to it.
2022-12-16 09:09:06 +04:00
justheuristic
91898c3c90
Switch to speedtest-cli (#157)
This pullrequest removes custom speed_test code in favour of speedtest-cli module.
This is necessary to ensure that random warnings / print-outs do not mess with our outputs.

Co-authored-by: Max Ryabinin <mryabinin0@gmail.com>
2022-12-15 15:21:33 +03:00
Max Ryabinin
34644f13e1
Downgrade CUDA in Docker image to 11.0.3 (#145)
* Downgrade CUDA in Docker image to 11.0.3

* Remove development deps from the image
2022-12-15 10:14:29 +03:00
Artem Chumachenko
7911c2641d
Update advanced notebooks (#148)
Update examples

Co-authored-by: Alexander Borzunov <borzunov.alexander@gmail.com>
2022-12-15 09:28:25 +04:00
Alexander Borzunov
668b736031
Fix logging: do not duplicate lines, enable colors in Colab (#156) 2022-12-15 09:12:18 +04:00
Alexander Borzunov
041ad20891
Check reachability automatically and give advice how to fix it (#155)
1. If we connect to the **public swarm**, the server now **automatically checks its DHT's reachability** from the outside world using API at http://health.petals.ml This is important to disallow unreachable servers to proceed (they create issues for the clients, such as repetitive retries).

    If http://health.petals.ml is down, the server proceeds without the check (so we don't depend on it). However, if health.petals.ml is up and explicitly tells us that we are unrechable, the server shows the reason of that and how to solve it.

    The check may be disabled with the `--skip_reachability_check` option (though I can't imagine cases where someone needs to use it).

2. Added `--port` and `--public_ip` as **simplified convenience options** for users not familiar with `--host_maddrs` and `--announce_maddrs`.
2022-12-15 05:04:09 +04:00
Alexander Borzunov
73df69a117
Reset MemoryCache during rebalancings (#154)
Before this PR, if there were open inference sessions right when rebalancing is triggered, their cache was never properly destroyed.
2022-12-15 00:11:46 +04:00
Max Ryabinin
bd91be27ea
Add missing methods for SamplingAlgorithm, fix docstrings (#107)
* Add missing methods for SamplingAlgorithm, fix docstrings

* Add SamplingAlgorithm to _choose_sample_algorithm

* Add test_sampling

* Add a warning if sampling options were passed, but do_sample=False

* Skip the sampling test for now

Co-authored-by: Alexander Borzunov <borzunov.alexander@gmail.com>
2022-12-13 20:09:15 +03:00
Max Ryabinin
a0e8bbd28d
Fix arguments in remove_old_models.py (#153)
* Fix arguments in remove_old_models.py

* Remove unnecessary args.author

* Fix the GitHub Action as well
2022-12-13 19:01:12 +03:00
Alexander Borzunov
701ec7e53e
Clean up disk space (#152) 2022-12-13 18:50:43 +04:00
justheuristic
b04982c1a2
Bump transformers to 4.25.1 (#151)
- latest accelerate, transformers, huggingface_hub
- rearrange attention caches to support https://github.com/huggingface/transformers/pull/18344
- remove unused code
- fix edge case where session crashes when receiving seq length 0
- assert transformer version when importing WrappedBloomBlock

Co-authored-by: Alexander Borzunov <borzunov.alexander@gmail.com>
Co-authored-by: Max Ryabinin <mryabinin0@gmail.com>
2022-12-13 11:03:49 +03:00
Alexander Borzunov
e4dc938dfe
Fix OOMs during server rebalancing (#150)
The cause of OOMs were the cyclic references `TransformerBackend <-> PrioritizedTaskPool` that could not have been garbage collected properly. Still, I've added explicit tensor removal just in case.
2022-12-13 00:44:40 +04:00
Alexander Borzunov
83d9493b6c
Improve block size calculations (#149) 2022-12-12 13:15:23 +04:00
Aleksandr Borzunov
f42e559c77 Update README.md 2022-12-09 17:54:00 +00:00
Alexander Borzunov
6beb686909
Add link to privacy & security Wiki (#144) 2022-12-09 21:40:29 +04:00
Alexander Borzunov
84fec81543
Suppress asyncio error logs by default (#142) 2022-12-09 04:32:26 +04:00
Alexander Borzunov
e99bf36647
Use common folder for all caches, make it a volume in Dockerfile (#141) 2022-12-09 03:54:57 +04:00
Alexander Borzunov
5f50ea9c79
Update Anaconda instructions (#140) 2022-12-09 01:16:33 +04:00
Alexander Borzunov
e1d8793f00
Show route on client (#139) 2022-12-08 20:26:33 +04:00
Alexander Borzunov
4cb0ac4718
Update texts in "Terms of use" and "Privacy and security" sections (#138) 2022-12-08 04:41:02 +04:00
Alexander Borzunov
a94c91d870
Add Docker commands, use permanent Discord links (#137) 2022-12-08 01:54:22 +04:00
Alexander Borzunov
77a00e17f0
Fix "could not unlink the shared memory file" during rebalancing (#135) 2022-12-07 12:59:34 +04:00
Alexander Borzunov
318d690a5c
Fix waiting until free memory is available (#136) 2022-12-07 02:29:54 +04:00
Alexander Borzunov
e8fac92e59
Allow .generate() to reuse existing inference session (#132) 2022-12-06 00:20:26 +04:00
Alexander Borzunov
1fe3716589
Don't ban servers in case of client-caused handler errors (#134) 2022-12-05 19:05:43 +04:00
Alexander Borzunov
66f1799d32
Set default --step_timeout to 5 min (#133) 2022-12-05 13:44:18 +04:00
Alexander Borzunov
b873d92ffa
Update README.md 2022-12-04 22:48:51 +04:00
Alexander Borzunov
5d5d2666b8
Mention parallel inference 2022-12-04 22:48:32 +04:00
Alexander Borzunov
955eae30b3
Mention 1 sec/token explicitly 2022-12-04 22:10:15 +04:00
Alexander Borzunov
33c210b973
Update Colab notebook 2022-12-03 20:38:18 +04:00
Alexander Borzunov
f56edaa13f
Fix inference and rpc_info() fault tolerance (#131) 2022-12-03 19:28:15 +04:00
justheuristic
79a4308992
Clear trigger before engaging in update (#130)
Update sequence_manager.py
2022-12-03 17:42:52 +03:00
Alexander Borzunov
b8e1c1b7f5
Revert to hivemind==1.1.3 for stability (#129) 2022-12-03 17:36:05 +04:00
justheuristic
68c85e7492
Avoid synchronous updates, ban peers based on request outcome (#127)
- sequence_manager now takes care for its own updated-ness - no need to manually update it
- if a peer fails a request, sequence manager will ban this peer temporarily. Ban times increase with failure streaks



Co-authored-by: Alexander Borzunov <borzunov.alexander@gmail.com>
2022-12-03 16:13:15 +03:00
Alexander Borzunov
9dbf5e2e6f
Set dht.num_workers = n_layer, update_period = 150, expiration = 300 (#125) 2022-12-03 15:26:57 +04:00
Max Ryabinin
3ca8b4f082
Fix typos with codespell (#126) 2022-12-03 14:19:37 +03:00
justheuristic
8491ed2bd3
Add checks for forward() inputs on the client side (#123) 2022-12-03 15:02:48 +04:00
Max Ryabinin
055f85b83e
Call block.load_state_dict only once (#124) 2022-12-03 15:01:56 +04:00
Artem Chumachenko
0855aa7347
Update notebooks to use full BLOOM-176B (#104)
Co-authored-by: Alexander Borzunov <borzunov.alexander@gmail.com>
2022-12-03 14:09:21 +04:00
Max Ryabinin
4ffb4d83c7
Remove "-r" when installing Petals in examples (#122) 2022-12-03 11:21:45 +04:00
Alexander Borzunov
d29ef70c85
Update README.md 2022-12-03 01:14:02 +04:00
Alexander Borzunov
1d9aa77697
Update README.md 2022-12-03 00:46:52 +04:00
Alexander Borzunov
da36470a4b
Update README.md 2022-12-03 00:46:08 +04:00
Alexander Borzunov
81b94df14b
Rework readme, move code example to the top, link draft of Colab (#118) 2022-12-03 00:17:57 +04:00
Alexander Borzunov
893987ebf8
Require hivemind==1.1.4 with p2pd v0.3.13 (#121) 2022-12-03 00:16:14 +04:00
Alexander Borzunov
fc6722576b
Choose --num_blocks for bigscience/bloom-petals automatically (#119) 2022-12-02 23:17:44 +04:00
Alexander Borzunov
f72c220404
Suppress quantization warning and fix dtype defaults in compute benchmark (#117) 2022-12-02 20:07:28 +04:00