awesome-llama: https://github.com/vllm-project/vllm
amd blackwell cuda deepseek deepseek-v3 gpt gpt-oss inference kimi llama llm llm-serving model-serving moe openai pytorch qwen qwen3 tpu transformer
Score: 34.95675711385041
Last synced: about 4 hours ago
JSON representation
Repository metadata:
A high-throughput and memory-efficient inference and serving engine for LLMs
- Host: GitHub
- URL: https://github.com/vllm-project/vllm
- Owner: vllm-project
- License: apache-2.0
- Created: 2023-02-09T11:23:20.000Z (over 3 years ago)
- Default Branch: main
- Last Pushed: 2026-08-11T21:35:06.000Z (1 day ago)
- Last Synced: 2026-08-11T21:50:48.712Z (1 day ago)
- Topics: amd, blackwell, cuda, deepseek, deepseek-v3, gpt, gpt-oss, inference, kimi, llama, llm, llm-serving, model-serving, moe, openai, pytorch, qwen, qwen3, tpu, transformer
- Language: Python
- Homepage: https://vllm.ai
- Size: 248 MB
- Stars: 88,791
- Watchers: 591
- Forks: 20,552
- Open Issues: 6,448
-
Metadata Files:
- Readme: README.md
- Contributing: CONTRIBUTING.md
- Funding: .github/FUNDING.yml
- License: LICENSE
- Code of conduct: CODE_OF_CONDUCT.md
- Codeowners: .github/CODEOWNERS
- Security: SECURITY.md
- Governance: docs/governance/collaboration.md
- Agents: AGENTS.md
- Claude: CLAUDE.md
- Dco: DCO
-
Funding:
- Github: vllm-project
- Open collective: vllm
Owner metadata:
- Name: vLLM
- Login: vllm-project
- Email:
- Kind: organization
- Description:
- Website:
- Location:
- Twitter:
- Company:
- Icon url: https://avatars.githubusercontent.com/u/136984999?v=4
- Repositories: 44
- Last Synced at: 2026-08-07T18:14:28.861Z
- Profile URL: https://github.com/vllm-project
GitHub Events
Total
- Commit comment event: 1
- Create event: 73
- Delete event: 130
- Fork event: 282
- Issue comment event: 4356
- Issues event: 584
- Pull request event: 812
- Pull request review comment event: 1372
- Pull request review event: 1717
- Push event: 661
- Release event: 1
- Watch event: 1388
- Total: 11377
Last Year
- Commit comment event: 1
- Create event: 73
- Delete event: 130
- Fork event: 282
- Issue comment event: 4356
- Issues event: 584
- Pull request event: 812
- Pull request review comment event: 1372
- Pull request review event: 1717
- Push event: 661
- Release event: 1
- Watch event: 1388
- Total: 11377
Committers metadata
Last synced: 2 days ago
Total Commits: 19,781
Total Committers: 2,897
Avg Commits per committer: 6.828
Development Distribution Score (DDS): 0.955
Commits in past year: 11,333
Committers in past year: 1,949
Avg Commits per committer in past year: 5.815
Development Distribution Score (DDS) in past year: 0.965
| Name | Commits | |
|---|---|---|
| Cyrus Leung | t****c@c****k | 900 |
| Woosuk Kwon | w****n@b****u | 832 |
| Michael Goin | m****4@g****m | 588 |
| Harry Mellor | 1****r | 546 |
| youkaichao | y****o@g****m | 474 |
| Nick Hill | n****l@r****m | 458 |
| Isotr0py | m****f@m****n | 429 |
| Wentao Ye | 4****6 | 395 |
| Jee Jee Li | p****e@g****m | 322 |
| Andreas Karatzas | a****a@a****m | 311 |
| Nicolò Lucchesi | n****s@r****m | 227 |
| Roger Wang | 1****6 | 222 |
| Lucas Wilkinson | L****n | 219 |
| Simon Mo | s****o@h****m | 196 |
| Russell Bryant | r****t@r****m | 182 |
| Kevin H. Luu | k****0@g****m | 172 |
| Chauncey | c****g@g****m | 170 |
| Reid | 6****1 | 160 |
| wang.yuqi | y****g@d****o | 154 |
| Tyler Michael Smith | t****r@n****m | 152 |
| Matthew Bonanni | m****i@r****m | 136 |
| Li, Jiang | j****i@i****m | 134 |
| Zhuohan Li | z****3@g****m | 127 |
| Kunshang Ji | k****i@i****m | 118 |
| Varun Sundar Rabindranath | v****8@g****m | 104 |
| Chen Zhang | z****9@o****m | 102 |
| bnellnm | 4****m | 101 |
| Robert Shaw | 1****t | 101 |
| Thomas Parnell | t****a@z****m | 91 |
| Ning Xie | a****g@g****m | 84 |
| and 2867 more... | ||
Issue and Pull Request metadata
Last synced: about 8 hours ago
Total issues: 9,349
Total pull requests: 17,106
Average time to close issues: 3 months
Average time to close pull requests: 12 days
Total issue authors: 5,382
Total pull request authors: 2,444
Average comments per issue: 2.64
Average comments per pull request: 3.07
Merged pull request: 8,624
Bot issues: 1
Bot pull requests: 36
Past year issues: 696
Past year pull requests: 2,081
Past year average time to close issues: 6 days
Past year average time to close pull requests: 4 days
Past year issue authors: 511
Past year pull request authors: 743
Past year average comments per issue: 1.19
Past year average comments per pull request: 2.0
Past year merged pull request: 788
Past year bot issues: 0
Past year bot pull requests: 4
Top Issue Authors
- simon-mo (84)
- WoosukKwon (68)
- youkaichao (67)
- robertgshaw2-redhat (48)
- mgoin (47)
- pseudotensor (47)
- DarkLight1337 (40)
- hahmad2008 (34)
- sleepwalker2017 (32)
- njhill (27)
- yewentao256 (26)
- DefTruth (25)
- cadedaniel (24)
- zhuohan123 (23)
- tdoublep (23)
Top Pull Request Authors
- DarkLight1337 (766)
- youkaichao (697)
- WoosukKwon (559)
- mgoin (559)
- Isotr0py (377)
- hmellor (350)
- jeejeelee (281)
- njhill (277)
- ywang96 (254)
- russellb (234)
- simon-mo (229)
- robertgshaw2-neuralmagic (188)
- tlrmchlsmth (185)
- reidliu41 (172)
- LucasWilkinson (172)
Top Issue Labels
- bug (4,199)
- stale (1,510)
- usage (1,274)
- feature request (1,174)
- installation (309)
- performance (283)
- misc (233)
- RFC (229)
- documentation (173)
- new model (164)
- good first issue (98)
- ci-failure (62)
- rocm (58)
- help wanted (46)
- ray (39)
- ready (35)
- unstale (32)
- new-model (29)
- ci/build (22)
- torch.compile (22)
- v1 (21)
- structured-output (17)
- tool-calling (14)
- P0 (13)
- duplicate (13)
- multi-modality (13)
- tpu (13)
- release (12)
- quantization (10)
- enhancement (8)
Top Pull Request Labels
- ready (7,213)
- v1 (2,037)
- ci/build (2,007)
- documentation (1,877)
- frontend (1,219)
- needs-rebase (721)
- tpu (446)
- rocm (431)
- multi-modality (413)
- speculative-decoding (327)
- bug (311)
- performance (283)
- structured-output (279)
- qwen (236)
- stale (180)
- tool-calling (180)
- llama (151)
- deepseek (143)
- new-model (140)
- quantization (102)
- gpt-oss (81)
- kv-connector (78)
- unstale (65)
- x86 CPU (54)
- nvidia (51)
- force-merge (48)
- perf-benchmarks (40)
- dependencies (37)
- k3 (28)
- kimi (27)
Package metadata
- Total packages: 35
-
Total downloads:
- pypi: 5,488,723 last-month
- conda: 110 total
- Total docker downloads: 16,137
- Total dependent packages: 46 (may contain duplicates)
- Total dependent repositories: 5 (may contain duplicates)
- Total versions: 362
- Total maintainers: 29
- Total advisories: 60
nixpkgs-unstable: vllm
High-throughput and memory-efficient inference and serving engine for LLMs
- Homepage: https://github.com/vllm-project/vllm
- Documentation: https://github.com/NixOS/nixpkgs/blob/nixos-unstable/pkgs/development/python-modules/vllm/default.nix#L585
- Licenses: Apache-2.0
- Latest release: 0.15.1 (published 5 months ago)
- Last Synced: 2026-08-12T01:18:16.786Z (1 day ago)
- Versions: 2
- Dependent Packages: 0
- Dependent Repositories: 0
-
Rankings:
- Dependent repos count: 0.0%
- Dependent packages count: 0.0%
- Average: 0.034%
- Forks count: 0.053%
- Stargazers count: 0.084%
- Maintainers (4)
pypi.org: vllm
A high-throughput and memory-efficient inference and serving engine for LLMs
- Homepage: https://github.com/vllm-project/vllm
- Documentation: https://docs.vllm.ai/en/latest/
- Licenses: Apache-2.0
- Latest release: 0.27.1 (published 2 days ago)
- Last Synced: 2026-08-12T05:23:27.470Z (1 day ago)
- Versions: 95
- Dependent Packages: 46
- Dependent Repositories: 5
- Downloads: 5,415,672 Last month
- Docker Downloads: 16,137
-
Rankings:
- Stargazers count: 0.327%
- Downloads: 1.489%
- Forks count: 1.559%
- Docker downloads count: 3.168%
- Average: 3.449%
- Dependent repos count: 6.776%
- Dependent packages count: 7.373%
- Maintainers (4)
-
Advisories:
- vLLM denial of service via prompt embeds on M-RoPE models
- vLLM: Speech-to-text upload size limit is enforced after full UploadFile read
- vLLM: ReDoS via structured_outputs.regex compiled without timeout in xgrammar and outlines backends
- vLLM has Remote DoS via Invalid Recovered Token Reinjection
- vLLM: Processing differential in multi-channel audio downmixing enables hidden-input/moderation bypass for audio models
- Duplicate Advisory: image EXIF Rotation & PNG tRNS Transparency Not Normalized, Causing Mismatch Between Model Input and Expectations
- vLLM: OOM Denial of Service via Audio Decompression Bomb
- vLLM: incomplete CVE-2026-22778 fix leaks PIL repr addresses via Anthropic router
- vLLM: GGUF dequantize kernel int truncation exposes uninitialized GPU memory in multi-tenant serving
- vLLM: image EXIF Rotation & PNG tRNS Transparency Not Normalized, Causing Mismatch Between Model Input and Expectations
- vLLM: temperature=NaN and temperature=Infinity bypass validation and propagate to GPU kernels
- vLLM: OpenAI auth bypass
- vLLM: Security Check Bypass via assert Statement in Activation Function Loading Allows Arbitrary Code Execution
- vLLM's Artifact Pin Decay allows pinned deployments to load unpinned code, weights, and processors
- vllm has Improper Resource Shutdown or Release
- vLLM: extract_hidden_states speculative decoding crashes server on any request with penalty parameters
- vLLM Vulnerable to Remote DoS via Special-Token Placeholders
- vLLM makes Use of Uninitialized Resource
- vLLM: Denial of Service via Unbounded Frame Count in video/jpeg Base64 Processing
- vLLM: Server-Side Request Forgery (SSRF) in `download_bytes_from_url `
- vLLM: Unauthenticated OOM Denial of Service via Unbounded `n` Parameter in OpenAI API Server
- vLLM has Hardcoded Trust Override in Model Files Enables RCE Despite Explicit User Opt-Out
- vLLM has SSRF Protection Bypass
- vLLM has RCE In Video Processing
- vLLM vulnerable to Server-Side Request Forgery (SSRF) through MediaConnector
- vLLM affected by RCE via auto_map dynamic module loading during model initialization
- vLLM is vulnerable to DoS in Idefics3 vision models via image payload with ambiguous dimensions
- vLLM introduced enhanced protection for CVE-2025-62164
- vLLM vulnerable to remote code execution via transformers_utils/get_config
- vLLM vulnerable to DoS via large Chat Completion or Tokenization requests with specially crafted `chat_template_kwargs`
- vLLM vulnerable to DoS with incorrect shape of multimodal embedding inputs
- vLLM deserialization vulnerability leading to DoS and potential RCE
- vLLM is vulnerable to Server-Side Request Forgery (SSRF) through `MediaConnector` class
- vLLM: Resource-Exhaustion (DoS) through Malicious Jinja Template in OpenAI-Compatible Server
- vLLM is vulnerable to timing attack at bearer auth
- vLLM has remote code execution vulnerability in the tool call parser for Qwen3-Coder
- vllm API endpoints vulnerable to Denial of Service Attacks
- vLLM Tool Schema allows DoS via Malformed pattern and type Fields
- vLLM allows clients to crash the openai server with invalid regex
- vLLM DOS: Remotely kill vllm over http with invalid JSON schema
- vLLM has a Weakness in MultiModalHasher Image Hashing Implementation
- Potential Timing Side-Channel Vulnerability in vLLM’s Chunk-Based Prefix Caching
- vLLM vulnerable to Regular Expression Denial of Service
- vLLM has a Regular Expression Denial of Service (ReDoS, Exponential Complexity) Vulnerability in `pythonic_tool_parser.py`
- vLLM Allows Remote Code Execution via PyNcclPipe Communication Service
- Remote Code Execution Vulnerability in vLLM Multi-Node Cluster Configuration
- vLLM: Quadratic Time Complexity in Input Token Processing leads to denial of service
- vLLM Vulnerable to Remote Code Execution via Mooncake Integration
- Data exposure via ZeroMQ on multi-node vLLM deployment
- CVE-2025-24357 Malicious model remote code execution fix bypass with PyTorch < 2.6.0
- vLLM vulnerable to Denial of Service by abusing xgrammar cache
- vLLM deserialization vulnerability in vllm.distributed.GroupCoordinator.recv_object
- vLLM allows Remote Code Execution by Pickle Deserialization via AsyncEngineRPCServer() RPC server entrypoints
- vLLM Deserialization of Untrusted Data vulnerability
- vLLM Allows Remote Code Execution via Mooncake Integration
- vLLM denial of service via outlines unbounded cache on disk
- vLLM uses Python 3.12 built-in hash() which leads to predictable hash collisions in prefix cache
- vllm: Malicious model to RCE by torch.load in hf_model_weights_iterator
- vLLM Denial of Service via the best_of parameter
- vLLM denial of service vulnerability
proxy.golang.org: github.com/vllm-project/vllm
- Homepage:
- Documentation: https://pkg.go.dev/github.com/vllm-project/vllm#section-documentation
- Licenses: apache-2.0
- Latest release: v0.27.1 (published 2 days ago)
- Last Synced: 2026-08-11T22:32:38.033Z (1 day ago)
- Versions: 83
- Dependent Packages: 0
- Dependent Repositories: 0
-
Rankings:
- Dependent packages count: 9.069%
- Average: 9.647%
- Dependent repos count: 10.226%
pypi.org: personal-test-vllm-tpu
A high-throughput and memory-efficient inference and serving engine for LLMs
- Homepage: https://github.com/vllm-project/vllm
- Documentation: https://docs.vllm.ai/en/latest/
- Licenses: apache-2.0
- Latest release: 0.10.1.1 (published 11 months ago)
- Last Synced: 2026-08-11T22:32:35.891Z (1 day ago)
- Versions: 1
- Dependent Packages: 0
- Dependent Repositories: 0
-
Rankings:
- Stargazers count: 0.09%
- Forks count: 0.12%
- Dependent packages count: 8.583%
- Average: 14.291%
- Dependent repos count: 48.372%
- Maintainers (1)
pypi.org: tmp-test-vllm-tpu
A high-throughput and memory-efficient inference and serving engine for LLMs
- Homepage: https://github.com/vllm-project/vllm
- Documentation: https://docs.vllm.ai/en/latest/
- Licenses: apache-2.0
- Latest release: 0.0.1 (published 11 months ago)
- Last Synced: 2026-08-11T22:32:35.845Z (1 day ago)
- Versions: 1
- Dependent Packages: 0
- Dependent Repositories: 0
- Downloads: 168 Last month
-
Rankings:
- Stargazers count: 0.09%
- Forks count: 0.119%
- Dependent packages count: 8.523%
- Average: 14.76%
- Downloads: 17.027%
- Dependent repos count: 48.039%
- Maintainers (1)
pypi.org: vllm-test-tpu
A high-throughput and memory-efficient inference and serving engine for LLMs
- Homepage: https://github.com/vllm-project/vllm
- Documentation: https://vllm.readthedocs.io/en/latest/
- Licenses: Apache-2.0
- Latest release: 0.9.1 (published about 1 year ago)
- Last Synced: 2026-08-12T01:36:41.139Z (1 day ago)
- Versions: 2
- Dependent Packages: 0
- Dependent Repositories: 0
- Downloads: 15 Last month
-
Rankings:
- Stargazers count: 0.136%
- Forks count: 0.166%
- Dependent packages count: 9.113%
- Average: 15.191%
- Dependent repos count: 51.347%
- Maintainers (1)
pypi.org: vllm-fixed
A high-throughput and memory-efficient inference and serving engine for LLMs
- Homepage: https://github.com/vllm-project/vllm
- Documentation: https://vllm.readthedocs.io/en/latest/
- Licenses: Apache-2.0
- Latest release: 1.0.0 (published 10 months ago)
- Last Synced: 2026-08-12T01:36:39.306Z (1 day ago)
- Versions: 2
- Dependent Packages: 0
- Dependent Repositories: 0
- Downloads: 11 Last month
-
Rankings:
- Stargazers count: 0.086%
- Forks count: 0.112%
- Dependent packages count: 8.46%
- Average: 15.642%
- Downloads: 21.732%
- Dependent repos count: 47.822%
- Maintainers (1)
pypi.org: vllm-cpu-nightly
A high-throughput and memory-efficient inference and serving engine for LLMs
- Homepage: https://github.com/vllm-project/vllm
- Documentation: https://docs.vllm.ai/en/latest/
- Licenses: Apache-2.0
- Latest release: 0.27.1.dev202608110727 (published 2 days ago)
- Last Synced: 2026-08-12T01:36:42.659Z (1 day ago)
- Versions: 70
- Dependent Packages: 0
- Dependent Repositories: 0
- Downloads: 14,903 Last month
-
Rankings:
- Dependent packages count: 7.058%
- Downloads: 14.097%
- Average: 20.361%
- Dependent repos count: 39.927%
- Maintainers (1)
pypi.org: vllm-hust
A high-throughput and memory-efficient inference and serving engine for LLMs
- Homepage: https://github.com/vllm-project/vllm
- Documentation: https://docs.vllm.ai/en/latest/
- Licenses: Apache-2.0
- Latest release: 0.17.2.post1 (published 4 months ago)
- Last Synced: 2026-08-12T01:36:41.562Z (1 day ago)
- Versions: 4
- Dependent Packages: 0
- Dependent Repositories: 0
- Downloads: 25 Last month
-
Rankings:
- Dependent packages count: 7.609%
- Average: 25.316%
- Dependent repos count: 43.024%
- Maintainers (1)
pypi.org: vllm-musa
vLLM platform plugin for Moore Threads MUSA GPUs
- Homepage: https://github.com/vllm-project/vllm
- Documentation: https://vllm-musa.readthedocs.io/
- Licenses: Apache-2.0
- Latest release: 0.1.1 (published 8 months ago)
- Last Synced: 2026-08-12T01:36:39.469Z (1 day ago)
- Versions: 3
- Dependent Packages: 0
- Dependent Repositories: 0
- Downloads: 53 Last month
-
Rankings:
- Dependent packages count: 8.356%
- Average: 27.794%
- Dependent repos count: 47.231%
- Maintainers (1)
pypi.org: vllm-usf
vLLM-USF: A high-throughput and memory-efficient inference engine for LLMs (USF Custom Build)
- Homepage: https://github.com/vllm-project/vllm
- Documentation: https://docs.vllm.ai
- Licenses: Apache-2.0
- Latest release: 0.0.2 (published 10 months ago)
- Last Synced: 2026-08-12T01:36:38.967Z (1 day ago)
- Versions: 2
- Dependent Packages: 0
- Dependent Repositories: 0
- Downloads: 25 Last month
-
Rankings:
- Dependent packages count: 8.449%
- Average: 28.103%
- Dependent repos count: 47.757%
- Maintainers (1)
anaconda.org: vllm
vLLM is a fast and easy-to-use library for LLM inference and serving.
- Homepage: https://vllm.ai/
- Licenses: Apache-2.0
- Latest release: 0.21.0 (published 21 days ago)
- Last Synced: 2026-07-23T09:43:58.382Z (21 days ago)
- Versions: 2
- Dependent Packages: 0
- Dependent Repositories: 0
- Downloads: 110 Total
-
Rankings:
- Stargazers count: 1.022%
- Forks count: 2.486%
- Average: 29.538%
- Dependent packages count: 36.615%
- Dependent repos count: 39.924%
- Downloads: 67.642%
pypi.org: vllm-tpu
A high-throughput and memory-efficient inference and serving engine for LLMs
- Homepage: https://github.com/vllm-project/vllm
- Documentation: https://docs.vllm.ai/en/latest/
- Licenses: Apache-2.0
- Latest release: 0.26.0 (published 13 days ago)
- Last Synced: 2026-08-12T01:36:41.162Z (1 day ago)
- Versions: 27
- Dependent Packages: 0
- Dependent Repositories: 0
- Downloads: 57,464 Last month
-
Rankings:
- Dependent packages count: 9.078%
- Average: 30.114%
- Dependent repos count: 51.149%
- Maintainers (2)
pypi.org: vllm-emissary
A high-throughput and memory-efficient inference and serving engine for LLMs
- Homepage: https://github.com/vllm-project/vllm
- Documentation: https://vllm.readthedocs.io/en/latest/
- Licenses: Apache Software License
- Latest release: 0.1.0 (published over 1 year ago)
- Last Synced: 2026-08-11T22:32:36.162Z (1 day ago)
- Versions: 2
- Dependent Packages: 0
- Dependent Repositories: 0
- Downloads: 17 Last month
-
Rankings:
- Dependent packages count: 9.343%
- Average: 30.984%
- Dependent repos count: 52.625%
- Maintainers (1)
pypi.org: ai-dynamo-vllm
A high-throughput and memory-efficient inference and serving engine for LLMs
- Homepage: https://github.com/vllm-project/vllm
- Documentation: https://vllm.readthedocs.io/en/latest/
- Licenses: Apache Software License
- Latest release: 0.8.4 (published over 1 year ago)
- Last Synced: 2026-08-11T22:32:36.145Z (1 day ago)
- Versions: 7
- Dependent Packages: 0
- Dependent Repositories: 0
- Downloads: 55 Last month
-
Rankings:
- Dependent packages count: 9.463%
- Average: 31.377%
- Dependent repos count: 53.292%
- Maintainers (1)
pypi.org: wxy-test
A high-throughput and memory-efficient inference and serving engine for LLMs
- Homepage: https://github.com/vllm-project/vllm
- Documentation: https://docs.vllm.ai/en/latest/
- Licenses: Apache-2.0
- Latest release: 0.19.0 (published 4 months ago)
- Last Synced: 2026-08-11T22:32:36.138Z (1 day ago)
- Versions: 3
- Dependent Packages: 0
- Dependent Repositories: 0
- Downloads: 14 Last month
-
Rankings:
- Dependent packages count: 9.652%
- Average: 31.999%
- Dependent repos count: 54.345%
- Maintainers (1)
pypi.org: vllm-npu
A high-throughput and memory-efficient inference and serving engine for LLMs
- Homepage: https://github.com/vllm-project/vllm
- Documentation: https://vllm.readthedocs.io/en/latest/
- Licenses: Apache 2.0
- Latest release: 0.4.2 (published over 1 year ago)
- Last Synced: 2026-08-12T01:36:41.234Z (1 day ago)
- Versions: 3
- Dependent Packages: 0
- Dependent Repositories: 0
- Downloads: 29 Last month
-
Rankings:
- Dependent packages count: 9.761%
- Average: 32.353%
- Dependent repos count: 54.946%
- Maintainers (1)
pypi.org: vllm-rocm
A high-throughput and memory-efficient inference and serving engine for LLMs with AMD GPU support
- Homepage: https://github.com/vllm-project/vllm
- Documentation: https://vllm.readthedocs.io/en/latest/
- Licenses: Apache 2.0
- Latest release: 0.6.3 (published almost 2 years ago)
- Last Synced: 2026-08-12T01:36:42.497Z (1 day ago)
- Versions: 1
- Dependent Packages: 0
- Dependent Repositories: 0
- Downloads: 65 Last month
-
Rankings:
- Dependent packages count: 10.195%
- Average: 33.788%
- Dependent repos count: 57.381%
- Maintainers (1)
pypi.org: vllm-acc
A high-throughput and memory-efficient inference and serving engine for LLMs
- Homepage: https://github.com/vllm-project/vllm
- Documentation: https://vllm.readthedocs.io/en/latest/
- Licenses: Apache 2.0
- Latest release: 0.4.1 (published over 2 years ago)
- Last Synced: 2026-08-12T01:36:33.823Z (1 day ago)
- Versions: 8
- Dependent Packages: 0
- Dependent Repositories: 0
- Downloads: 53 Last month
-
Rankings:
- Dependent packages count: 9.449%
- Average: 35.896%
- Dependent repos count: 62.342%
- Maintainers (1)
pypi.org: vllm-online
A high-throughput and memory-efficient inference and serving engine for LLMs
- Homepage: https://github.com/vllm-project/vllm
- Documentation: https://vllm.readthedocs.io/en/latest/
- Licenses: Apache 2.0
- Latest release: 0.4.2 (published over 2 years ago)
- Last Synced: 2026-08-12T01:36:33.075Z (1 day ago)
- Versions: 2
- Dependent Packages: 0
- Dependent Repositories: 0
- Downloads: 17 Last month
-
Rankings:
- Dependent packages count: 9.472%
- Average: 35.984%
- Dependent repos count: 62.495%
- Maintainers (1)
pypi.org: tilearn-infer
A high-throughput and memory-efficient inference and serving engine for LLMs
- Homepage: https://github.com/vllm-project/vllm
- Documentation: https://vllm.readthedocs.io/en/latest/
- Licenses: Apache 2.0
- Latest release: 0.3.3 (published over 2 years ago)
- Last Synced: 2026-08-12T01:36:33.540Z (1 day ago)
- Versions: 3
- Dependent Packages: 0
- Dependent Repositories: 0
- Downloads: 19 Last month
-
Rankings:
- Dependent packages count: 9.569%
- Average: 36.35%
- Dependent repos count: 63.131%
- Maintainers (1)
pypi.org: tilearn-test01
A high-throughput and memory-efficient inference and serving engine for LLMs
- Homepage: https://github.com/vllm-project/vllm
- Documentation: https://vllm.readthedocs.io/en/latest/
- Licenses: Apache 2.0
- Latest release: 0.1 (published over 2 years ago)
- Last Synced: 2026-08-12T01:36:33.766Z (1 day ago)
- Versions: 1
- Dependent Packages: 0
- Dependent Repositories: 0
- Downloads: 14 Last month
-
Rankings:
- Dependent packages count: 9.585%
- Average: 36.409%
- Dependent repos count: 63.233%
- Maintainers (1)
pypi.org: vllm-xft
A high-throughput and memory-efficient inference and serving engine for LLMs
- Homepage: https://github.com/vllm-project/vllm
- Documentation: https://vllm.readthedocs.io/en/latest/
- Licenses: Apache 2.0
- Latest release: 0.5.5.4 (published over 1 year ago)
- Last Synced: 2026-08-12T01:36:34.569Z (1 day ago)
- Versions: 12
- Dependent Packages: 0
- Dependent Repositories: 0
- Downloads: 31 Last month
-
Rankings:
- Dependent packages count: 9.629%
- Average: 36.582%
- Dependent repos count: 63.535%
- Maintainers (2)
pypi.org: hive-vllm
a
- Homepage: https://github.com/vllm-project/vllm
- Documentation: https://vllm.readthedocs.io/en/latest/
- Licenses: Apache 2.0
- Latest release: 0.0.1 (published over 2 years ago)
- Last Synced: 2026-08-12T01:36:34.080Z (1 day ago)
- Versions: 1
- Dependent Packages: 0
- Dependent Repositories: 0
- Downloads: 20 Last month
-
Rankings:
- Dependent packages count: 9.784%
- Average: 37.172%
- Dependent repos count: 64.559%
- Maintainers (1)
pypi.org: nextai-vllm
A high-throughput and memory-efficient inference and serving engine for LLMs
- Homepage: https://github.com/vllm-project/vllm
- Documentation: https://vllm.readthedocs.io/en/latest/
- Licenses: Apache 2.0
- Latest release: 0.0.7 (published over 2 years ago)
- Last Synced: 2026-08-11T22:32:36.137Z (1 day ago)
- Versions: 6
- Dependent Packages: 0
- Dependent Repositories: 0
- Downloads: 20 Last month
-
Rankings:
- Dependent packages count: 9.941%
- Average: 37.779%
- Dependent repos count: 65.617%
- Maintainers (1)
pypi.org: vllm-consul
A high-throughput and memory-efficient inference and serving engine for LLMs
- Homepage: https://github.com/vllm-project/vllm
- Documentation: https://vllm.readthedocs.io/en/latest/
- Licenses: Apache 2.0
- Latest release: 0.2.1 (published almost 3 years ago)
- Last Synced: 2026-08-12T01:36:35.690Z (1 day ago)
- Versions: 5
- Dependent Packages: 0
- Dependent Repositories: 0
- Downloads: 33 Last month
-
Rankings:
- Dependent packages count: 9.355%
- Average: 38.747%
- Dependent repos count: 68.139%
- Maintainers (1)
artifacthub.io: kubeblocks/vllm-chat
Self-hosted chat/agent LLM on vLLM for ApeMind private deployment ("instance A" in the MinerU-3.x + vLLM design). Serves an OpenAI-compatible endpoint with tool-calling enabled. Default model pin: Qwen/Qwen3.6-27B-FP8 (single 80G card recommended; 40G minimum per vLLM official recipes). The MinerU document-parsing VLM ("instance B") is NOT this chart — enable deploy/mineru's openaiServer service for that.
- Homepage: https://docs.vllm.ai/
- Documentation: https://artifacthub.io/packages/helm/kubeblocks/vllm-chat
- Licenses: Unknown
- Latest release: 0.1.0 (published 3 days ago)
- Last Synced: 2026-08-12T01:36:41.006Z (1 day ago)
- Versions: 1
- Dependent Packages: 0
- Dependent Repositories: 0
- Downloads: 0 Total
-
Rankings:
- Downloads: 0.0%
- Dependent repos count: 0.0%
- Dependent packages count: 0.0%
- Average: 100%
artifacthub.io: kubeblocks/vllm-embedding
Self-hosted multimodal embedding service on vLLM for ApeMind private deployment. Serves an OpenAI-compatible /v1/embeddings endpoint (vLLM --runner pooling). Default model pin: Qwen/Qwen3-VL-Embedding-2B (2026-01, Apache-2.0, SOTA multimodal embeddings, MRL dims 64..2048, ~6G VRAM — fits beside other services on a 24G card). Chat/agent LLM is deploy/vllm-chat; the MinerU document-parsing VLM is deploy/mineru openaiServer.
- Homepage: https://docs.vllm.ai/
- Documentation: https://artifacthub.io/packages/helm/kubeblocks/vllm-embedding
- Licenses: Unknown
- Latest release: 0.1.0 (published 3 days ago)
- Last Synced: 2026-08-12T01:36:42.008Z (1 day ago)
- Versions: 1
- Dependent Packages: 0
- Dependent Repositories: 0
- Downloads: 0 Total
-
Rankings:
- Downloads: 0.0%
- Dependent repos count: 0.0%
- Dependent packages count: 0.0%
- Average: 100%
nixpkgs-24.11: python311Packages.vllm
High-throughput and memory-efficient inference and serving engine for LLMs
- Homepage: https://github.com/vllm-project/vllm
- Documentation: https://github.com/NixOS/nixpkgs/blob/nixos-24.11/pkgs/development/python-modules/vllm/default.nix#L171
- Licenses: Apache-2.0
- Latest release: 0.6.2 (published 6 months ago)
- Last Synced: 2026-03-07T13:13:39.225Z (5 months ago)
- Versions: 1
- Dependent Packages: 0
- Dependent Repositories: 0
-
Rankings:
- Dependent repos count: 0.0%
- Dependent packages count: 0.0%
- Average: 100%
- Maintainers (2)
nixpkgs-unstable: python313Packages.vllm
High-throughput and memory-efficient inference and serving engine for LLMs
- Homepage: https://github.com/vllm-project/vllm
- Documentation: https://github.com/NixOS/nixpkgs/blob/nixos-unstable/pkgs/development/python-modules/vllm/default.nix#L585
- Licenses: Apache-2.0
- Latest release: 0.15.1 (published 5 months ago)
- Last Synced: 2026-05-14T17:14:40.043Z (3 months ago)
- Versions: 2
- Dependent Packages: 0
- Dependent Repositories: 0
-
Rankings:
- Dependent repos count: 0.0%
- Dependent packages count: 0.0%
- Average: 100%
- Maintainers (4)
nixpkgs-24.11: python312Packages.vllm
High-throughput and memory-efficient inference and serving engine for LLMs
- Homepage: https://github.com/vllm-project/vllm
- Documentation: https://github.com/NixOS/nixpkgs/blob/nixos-24.11/pkgs/development/python-modules/vllm/default.nix#L171
- Licenses: Apache-2.0
- Latest release: 0.6.2 (published 6 months ago)
- Last Synced: 2026-04-10T14:01:38.971Z (4 months ago)
- Versions: 1
- Dependent Packages: 0
- Dependent Repositories: 0
-
Rankings:
- Dependent repos count: 0.0%
- Dependent packages count: 0.0%
- Average: 100%
- Maintainers (2)
nixpkgs-24.05: python311Packages.vllm
A high-throughput and memory-efficient inference and serving engine for LLMs
- Homepage: https://github.com/vllm-project/vllm
- Documentation: https://github.com/NixOS/nixpkgs/blob/nixos-24.05/pkgs/development/python-modules/vllm/default.nix#L148
- Licenses: Apache-2.0
- Latest release: 0.3.3 (published 6 months ago)
- Last Synced: 2026-05-13T01:07:46.485Z (3 months ago)
- Versions: 1
- Dependent Packages: 0
- Dependent Repositories: 0
-
Rankings:
- Dependent repos count: 0.0%
- Dependent packages count: 0.0%
- Average: 100%
- Maintainers (2)
nixpkgs-24.05: python312Packages.vllm
A high-throughput and memory-efficient inference and serving engine for LLMs
- Homepage: https://github.com/vllm-project/vllm
- Documentation: https://github.com/NixOS/nixpkgs/blob/nixos-24.05/pkgs/development/python-modules/vllm/default.nix#L148
- Licenses: Apache-2.0
- Latest release: 0.3.3 (published 6 months ago)
- Last Synced: 2026-03-09T05:10:24.804Z (5 months ago)
- Versions: 1
- Dependent Packages: 0
- Dependent Repositories: 0
-
Rankings:
- Dependent repos count: 0.0%
- Dependent packages count: 0.0%
- Average: 100%
- Maintainers (2)
nixpkgs-unstable: python314Packages.vllm
High-throughput and memory-efficient inference and serving engine for LLMs
- Homepage: https://github.com/vllm-project/vllm
- Documentation: https://github.com/NixOS/nixpkgs/blob/nixos-unstable/pkgs/development/python-modules/vllm/default.nix#L585
- Licenses: Apache-2.0
- Latest release: 0.15.1 (published 5 months ago)
- Last Synced: 2026-03-05T16:21:08.613Z (5 months ago)
- Versions: 2
- Dependent Packages: 0
- Dependent Repositories: 0
-
Rankings:
- Maintainers (4)
artifacthub.io: infracloud-charts/vllm
A Helm chart for deploying models with vLLM
- Homepage:
- Documentation: https://artifacthub.io/packages/helm/infracloud-charts/vllm
- Licenses: Unknown
- Latest release: 0.2.2 (published over 1 year ago)
- Last Synced: 2026-08-12T01:36:43.816Z (1 day ago)
- Versions: 4
- Dependent Packages: 0
- Dependent Repositories: 0
- Downloads: 0 Total
-
Rankings:
- Downloads: 0.0%
- Dependent repos count: 0.0%
- Dependent packages count: 0.0%
- Average: 100%
Dependencies
- actions/checkout v2 composite
- actions/setup-python v2 composite
- actions/checkout v2 composite
- actions/setup-python v2 composite
- sphinx ==6.2.1
- sphinx-book-theme ==1.0.1
- sphinx-copybutton ==0.5.2
- mypy ==0.991 development
- pylint ==2.8.2 development
- pytest * development
- types-PyYAML * development
- types-requests * development
- types-setuptools * development
- yapf ==0.32.0 development
- fastapi *
- fschat *
- ninja *
- numpy *
- psutil *
- pydantic *
- ray *
- sentencepiece *
- torch >=2.0.0
- transformers >=4.28.0
- uvicorn *
- xformers >=0.0.19
- actions/checkout v3 composite
- actions/github-script v6 composite
- actions/setup-python v4 composite
- actions/upload-release-asset v1 composite