Collections: awesome-llama
https://github.com/vllm-project/vllm
A high-throughput and memory-efficient inference and serving engine for LLMs
amd blackwell cuda deepseek deepseek-v3 gpt gpt-oss inference kimi llama llm llm-serving model-serving moe openai pytorch qwen qwen3 tpu transformer
Last synced: 20 Aug 2026
https://github.com/huggingface/text-generation-inference
Large Language Model Text Generation Inference
bloom deep-learning falcon gpt inference nlp pytorch starcoder transformer
Last synced: 21 Aug 2026
https://github.com/xorbitsai/inference
Swap GPT for any LLM by changing a single line of code. Xinference lets you run open-source, speech, and multimodal models on cloud, on-prem, or your laptop — all through one unified, production-ready inference API.
artificial-intelligence chatglm deployment flan-t5 gemma ggml glm4 inference llama llama3 llamacpp llm machine-learning mistral openai-api pytorch qwen vllm whisper wizardlm
Last synced: 20 Aug 2026
https://github.com/ludwig-ai/ludwig
Low-code framework for building custom LLMs, neural networks, and other AI models
computer-vision data-centric data-science deep deep-learning deeplearning fine-tuning learning llama llama2 llm llm-training machine-learning machinelearning mistral ml natural-language natural-language-processing neural-network pytorch
Last synced: 21 Aug 2026
https://github.com/explosion/curated-transformers
🤖 A PyTorch library of curated Transformer models and their composable components
albert bert camembert dolly2 falcon gptneox llama llm llms nlp pytorch roberta transformer transformers xlm-roberta
Last synced: 21 Aug 2026
https://github.com/meta-llama/llama-cookbook
Welcome to the Llama Cookbook! This is your go to guide for Building with Llama: Getting started with Inference, Fine-Tuning, RAG. We also show you how to solve end to end problems using Llama model family and using them on various provider services
ai finetuning langchain llama llama2 llm machine-learning python pytorch vllm
Last synced: 21 Aug 2026
https://github.com/predibase/lorax
Multi-LoRA inference server that scales to 1000s of fine-tuned LLMs
fine-tuning gpt llama llm llm-inference llm-serving llmops lora model-serving pytorch transformers
Last synced: 21 Aug 2026
https://github.com/bigscience-workshop/petals
🌸 Run LLMs at home, BitTorrent-style. Fine-tuning and inference up to 10x faster than offloading
bloom chatbot deep-learning distributed-systems falcon gpt guanaco language-models large-language-models llama machine-learning mixtral neural-networks nlp pipeline-parallelism pretrained-models pytorch tensor-parallelism transformer volunteer-computing
Last synced: 21 Aug 2026
https://github.com/kyegomez/zeta
Build high-performance AI models with modular building blocks
attention-mechanism attention-model chatgpt ffns llms lucidrains openai pytorch pytorch-implementation pytorch-tutorial tensorflow transformer-architecture transformers
Last synced: 20 Aug 2026
https://github.com/tongjilibo/bert4torch
An elegent pytorch implement of transformers
belle bert bert4keras bert4torch chatglm large-language-models llama llm named-entity-recognition nlp pytorch relation-extraction seq2seq text-classification transformers
Last synced: 21 Aug 2026
https://github.com/Tongjilibo/bert4torch
An elegent pytorch implement of transformers
belle bert bert4keras bert4torch chatglm large-language-models llama llm named-entity-recognition nlp pytorch relation-extraction seq2seq text-classification transformers
Last synced: 21 Aug 2026
https://github.com/PhoebusSi/Alpaca-CoT
We unified the interfaces of instruction-tuning data (e.g., CoT data), multiple LLMs and parameter-efficient methods (e.g., lora, p-tuning) together for easy use. We welcome open-source enthusiasts to initiate any meaningful PR on this repo and integrate as many LLM related technologies as possible. 我们打造了方便研究人员上手和使用大模型等微调平台,我们欢迎开源爱好者发起任何有意义的pr!
alpaca chatglm chatgpt cot instruction-tuning llama llm lora moss p-tuning parameter-efficient pytorch tabul tabular-data tabular-model
Last synced: 20 Aug 2026
https://github.com/hiyouga/fastedit
🩹Editing large language models within 10 seconds⚡
bloom chatbots chatgpt falcon gpt large-language-models llama llms pytorch transformers
Last synced: 20 Aug 2026
https://github.com/hiyouga/FastEdit
🩹Editing large language models within 10 seconds⚡
bloom chatbots chatgpt falcon gpt large-language-models llama llms pytorch transformers
Last synced: 20 Aug 2026
https://github.com/alibaba/Megatron-LLaMA
Best practice for training LLaMA models in Megatron-LM
deepspeed distributed-training llama llm megatron-lm pretraining pytorch
Last synced: 21 Aug 2026
https://github.com/X-PLUG/mPLUG-Owl
mPLUG-Owl: The Powerful Multi-modal Large Language Model Family
alpaca chatbot chatgpt damo dialogue gpt gpt4 gpt4-api huggingface instruction-tuning large-language-models llama mplug mplug-owl multimodal pretraining pytorch transformer video visual-recognition
Last synced: 21 Aug 2026
https://github.com/BobaZooba/xllm
🦖 X—LLM: Cutting Edge & Easy LLM Finetuning
alpaca bitsandbytes cerebras chatgpt deep-learning deep-neural-networks gpt gpt-4 gptq large-language-models llama llama2 llm mistral openai pytorch torch vicuna zephyr
Last synced: 20 Aug 2026
https://github.com/bobazooba/xllm
🦖 X—LLM: Cutting Edge & Easy LLM Finetuning
alpaca bitsandbytes cerebras chatgpt deep-learning deep-neural-networks gpt gpt-4 gptq large-language-models llama llama2 llm mistral openai pytorch torch vicuna zephyr
Last synced: 20 Aug 2026
https://github.com/LinkSoul-AI/Chinese-Llama-2-7b
开源社区第一个能下载、能运行的中文 LLaMA2 模型!
deep-learning llama2 llama2-docker llm pytorch
Last synced: 20 Aug 2026
https://github.com/yuanzhoulvpi2017/zero_nlp
中文nlp解决方案(大模型、数据、模型、训练、推理)
bert chatglm-6b clip gpt gpt2 huggingface-transformers llama llama2 llava nlp pytorch text-generation transformers
Last synced: 21 Aug 2026
https://github.com/lxe/simple-llm-finetuner
Simple UI for LLM Model Finetuning
ai gpt-2 gpt-3 huggingface huggingface-transformers llama llm peft pytorch
Last synced: 21 Aug 2026
https://github.com/microsoft/sarathi-serve
A low-latency & high-throughput serving engine for LLMs
llama llm-inference pytorch transformer
Last synced: 20 Aug 2026
https://github.com/RahulSChand/gpu_poor
Calculate token/s & GPU memory requirement for any LLM. Supports llama.cpp/ggml/bnb/QLoRA quantization
ggml gpu huggingface language-model llama llama2 llamacpp llm pytorch quantization
Last synced: 21 Aug 2026
https://github.com/ashishpatel26/LLM-Finetuning
LLM Finetuning with peft
falcon fine-tuning huggingface llama llama2 llm llms lora peft pytorch text-generation
Last synced: 21 Aug 2026
https://github.com/minggnim/nlp-models
A repository for training transformer based models
chatbot chatbots ctransformers deeplearning falcon fine-tuning gpt-2 langchain llama2 llms multi-label-classification multi-task-learning nlp pytorch qdrant-vector-database transformers
Last synced: 20 Aug 2026
https://github.com/AI-Hypercomputer/jetstream-pytorch
PyTorch/XLA integration with JetStream (https://github.com/google/JetStream) for LLM inference"
attention batching gemma inference llama llama2 llm llm-inference model-serving pytorch tpu
Last synced: 21 Aug 2026
https://github.com/interestingLSY/swiftLLM
A tiny yet powerful LLM inference system tailored for researching purpose. vLLM-equivalent performance with only 2k lines of code (2% of vLLM).
cuda gpt inference inference-engine llama llm llm-inference llm-serving llmops mlops model-serving pytorch transformer transformers
Last synced: 21 Aug 2026
https://github.com/rohan-paul/LLM-FineTuning-Large-Language-Models
LLM (Large Language Model) FineTuning
gpt-3 gpt3-turbo large-language-models llama2 llm llm-finetuning llm-inference llm-serving llm-training mistral-7b open-source-llm pytorch
Last synced: 20 Aug 2026
https://github.com/stanleylsx/llms_tool
一个基于HuggingFace开发的大语言模型训练、测试工具。支持各模型的webui、终端预测,低参数量及全参数模型训练(预训练、SFT、RM、PPO、DPO)和融合、量化。
aquila aquila2 baichuan baichuan2 bloom chatglm chatglm2 chatglm3 deepspeed falcon internlm llama llama2 mistral moss pytorch qwen xverse
Last synced: 20 Aug 2026
https://github.com/mrdbourke/mac-ml-speed-test
A few quick scripts focused on testing TensorFlow/PyTorch/Llama 2 on macOS.
apple-silicon benchmark deep-learning llama2 llamacpp llm m1 m1-mac m2-mac m3-mac machine-learning macos metal metal-performance-shaders ml mps pytorch speedtest tensorflow2
Last synced: 20 Aug 2026
https://github.com/coderonion/awesome-cuda-and-hpc
🚀🚀🚀 This repository lists some awesome public CUDA, cuda-python, cuBLAS, cuDNN, CUTLASS, TensorRT, TensorRT-LLM, Triton, TVM, MLIR, PTX and High Performance Computing (HPC) projects.
awesome blas cublas cuda cudnn cutlass deepseek gemm gpu hpc llama llm mlir ptx pytorch tensorrt tensorrt-llm triton tvm vlm
Last synced: 20 Aug 2026
https://github.com/shm007g/LLaMA-Cult-and-More
Large Language Models for All, 🦙 Cult and More, Stay in touch !
alpaca chatgpt deepspeed ggml gpt gpt4 gptq llama llm loralib pytorch tensorflow transformers vicuna
Last synced: 20 Aug 2026
https://github.com/hkproj/pytorch-llama
LLaMA 2 implemented from scratch in PyTorch
language-model llama llama2 paper-implementations pytorch
Last synced: 21 Aug 2026
https://github.com/Gunale0926/SORSA
SORSA: Singular Values and Orthonormal Regularized Singular Vectors Adaptation of Large Language Models
deep-learning fine-tuning llama lora machine-learning nlp peft python pytorch rwkv sorsa svd transformer
Last synced: 20 Aug 2026
https://github.com/luchangli03/export_llama_to_onnx
export llama to onnx
llama llm llm-inference onnx onnxruntime pytorch
Last synced: 20 Aug 2026
https://github.com/nrimsky/CAA
Steering Llama 2 with Contrastive Activation Addition
llama2 pytorch
Last synced: 20 Aug 2026
https://github.com/jackaduma/Vicuna-LoRA-RLHF-PyTorch
A full pipeline to finetune Vicuna LLM with LoRA and RLHF on consumer hardware. Implementation of RLHF (Reinforcement Learning with Human Feedback) on top of the Vicuna architecture. Basically ChatGPT but with Vicuna
chatgpt finetune gpt llama llm lora peft ppo pytorch reward-models rlhf vicuna vicuna-7b
Last synced: 20 Aug 2026
https://github.com/YeonwooSung/ai_book
AI book for everyone
ai ai-ml cv deep-learning gpt-4 knowledge-distillation llama llm llmops machine-learning machinelearning mlops mlops-workflow nlp pytorch sagemaker tutorial xai
Last synced: 21 Aug 2026
https://github.com/akanyaani/miniLLAMA
A simplified LLAMA implementation for training and inference tasks.
llama llama-models llama2 llm minillama pytorch simple-llama
Last synced: 20 Aug 2026
https://github.com/A-baoYang/alpaca-7b-chinese
Finetune LLaMA-7B with Chinese instruction datasets
alpaca chatgpt deep-learning fine-tuning instruction-following llm lora nlp pytorch
Last synced: 20 Aug 2026
https://github.com/jackaduma/ChatGLM-LoRA-RLHF-PyTorch
A full pipeline to finetune ChatGLM LLM with LoRA and RLHF on consumer hardware. Implementation of RLHF (Reinforcement Learning with Human Feedback) on top of the ChatGLM architecture. Basically ChatGPT but with ChatGLM
chatglm chatglm-6b chatgpt deepspeed finetune gpt llama llm lora peft ppo pytorch reward-models rlhf
Last synced: 20 Aug 2026
https://github.com/sh0416/llama-classification
Text classification with Foundation Language Model LLaMA
classification foundation-model gpt language-model llama pytorch
Last synced: 20 Aug 2026
https://github.com/CoinCheung/gdGPT
Train llm (bloom, llama, baichuan2-7b, chatglm3-6b) with deepspeed pipeline mode. Faster than zero/zero++/fsdp.
baichuan2-7b bloom chatglm3-6b deepspeed flash-attention full-finetune llama2 llm mixtral-8x7b model-parallization nlp pipeline pytorch
Last synced: 20 Aug 2026
https://github.com/coderonion/awesome-mojo-max-mlir
A collection of some awesome public MAX platform, Mojo programming language and Multi-Level IR Compiler Framework(MLIR) projects.
awesome basalt chris-lattner cpp cuda large-language-models llama llama3 llm llvm machine-learning mlir modular mojo numpy pytorch rust tensor yolo yolov10
Last synced: 20 Aug 2026
https://github.com/ModelTC/QLLM
[ICLR 2024] This is the official PyTorch implementation of "QLLM: Accurate and Efficient Low-Bitwidth Quantization for Large Language Models"
llama llama2 llm post-training-quantization pytorch quantization transformers
Last synced: 20 Aug 2026
https://github.com/knagrecha/saturn
Saturn accelerates the training of large-scale deep learning models with a novel joint optimization approach.
ai ai-systems deep-learning distributed-systems distributed-training gpt gpu gpu-computing joint-optimization large-language-models large-model llama llama2 llmops mlops pytorch query-optimization systems-for-machine-learning training
Last synced: 21 Aug 2026
https://github.com/clabrugere/scratch-llm
From-scratch Llama 2-inspired Transformer in PyTorch for exploring tokenization, pre-training, and inference.
large-language-models llama llm pytorch
Last synced: 20 Aug 2026
https://github.com/b0kch01/llama-cpu
🦙 Inference code for LLaMA models (modified for cpu)
cpu llama meta pytorch
Last synced: 20 Aug 2026
https://github.com/Esmail-ibraheem/Axon
AI research lab🔬: implementations of AI papers and theoretical research: InstructGPT, llama, transformers, diffusion models, RLHF, etc...
arxiv-papers llama llms paper-implementations pytorch research-paper rlhf transformers
Last synced: 20 Aug 2026
https://github.com/jackaduma/Alpaca-LoRA-RLHF-PyTorch
A full pipeline to finetune Alpaca LLM with LoRA and RLHF on consumer hardware. Implementation of RLHF (Reinforcement Learning with Human Feedback) on top of the Alpaca architecture. Basically ChatGPT but with Alpaca
alpaca chatgpt deepspeed finetune gpt llama llm lora peft ppo pytorch reward-models rlhf
Last synced: 20 Aug 2026
https://github.com/mapluisch/LLaVA-CLI-with-multiple-images
LLaVA inference with multiple images at once for cross-image analysis.
image-concatenation image-processing inference llama2 llama2-13b llava lmm lmms pillow python python3 pytorch visual-question-answering vqa
Last synced: 20 Aug 2026
https://github.com/lxe/llama-peft-tuner
Tune LLaMa-7B on Alpaca Dataset using PEFT / LORA Based on @zphang's https://github.com/zphang/minimal-llama scripts.
llama llm lora machine-learning pytorch
Last synced: 20 Aug 2026
https://github.com/nouralmulhem/Enhancing-LLM-with-Jenkins-Knowledge
🚀 this project aims to develop an app using an existing open-source LLM with data collected for domain-specific Jenkins knowledge that can be fine-tuned locally and set up with a proper UI for the user to interact with.
chatbot llama2 llm python pytorch rag
Last synced: 20 Aug 2026
https://github.com/MoFHeka/LLaMA-Megatron
A LLaMA1/LLaMA12 Megatron implement.
llama llama2 llm llm-training megatron megatron-lm pytorch
Last synced: 20 Aug 2026
https://github.com/DominikLindorfer/pyAlpaca
Instruction-following LLaMA Model Trained with Deepspeed to Output Python-Code from General Instructions
alpaca deepspeed finetuning gpt gpt-3 gpt4 llama llm pyalpaca pytorch stanford-alpaca
Last synced: 20 Aug 2026
https://github.com/sytelus/nanuGPT
Simple, reliable and well tested training code for quick experiments with transformer based models
ai gpt llama llama2 llm machine-learning ml pytorch tokenization transformer
Last synced: 20 Aug 2026
https://github.com/GURPREETKAURJETHRA/Meta-LLAMA3-GenAI-UseCases-End-To-End-Implementation-Guides
META LLAMA3 GENAI Real World UseCases End To End Implementation Guide
chromadb fine-tuning fsdp generativeai huggingface langchain-python llama3 llama3-70b-8192 llama3-finetune llama3-meta-ai llama3-prompts llama3-rag ollama prompt-tuning pytorch qlora rag sagemaker streamlit
Last synced: 20 Aug 2026
https://github.com/soumyadip1995/BabyGPT
Something in the middle of Karpathy's mingpt model and video lectures, BabyGPT is an easy to use model on a much smaller scale (16 and 256 out channels , 5 heads, fine tuned).
artificial-intelligence gpt language-model large-language-models llama llama2 llm machine-learning natural-language-processing nlp python pytorch quantization transformer
Last synced: 20 Aug 2026
https://github.com/unipr-org/AI
AI - Intelligenza Artificiale presso l'Università degli Studi di Parma (6 CFU).
ai artificial-intelligence large-language-models llama llm neural-networks prolog pytorch
Last synced: 20 Aug 2026
https://github.com/cambridgeltl/prompt4bli
On Bilingual Lexicon Induction with Large Language Models (EMNLP 2023). Keywords: Bilingual Lexicon Induction, Word Translation, Large Language Models, LLMs.
bilingual-dictionary-induction bilingual-lexicon-extraction bilingual-lexicon-induction few-shot-learning in-context-learning large-language-models llama llms low-resource-machine-translation machine-translation mt5 multilingual-models multilingual-nlp prompt prompt-engineering prompting prompts pytorch word-translation zero-shot-learning
Last synced: 20 Aug 2026
https://github.com/reshalfahsi/image-captioning-mobilenet-llama3
Image Captioning With MobileNet-LLaMA 3
cnn flickr8k-dataset grouped-query-attention image-captioning image-text kv-cache llama3 mobilenetv3 nlp pytorch pytorch-lightning rms-norm rotary-position-embedding transformer
Last synced: 20 Aug 2026
https://github.com/Esmail-ibraheem/Tinyllamas-pytorch
Tinyllamas🦙 is an Extensible advanced language model framework, inspired by the original Llama model.
attention-mechanism llama llama2 llms paper-implementations pytorch
Last synced: 20 Aug 2026
https://github.com/ColinWu0403/LLaMA-2-hf-Chatbot
Chatbot from pretrained LLaMA-2 LLM model, fine-tuned with medical research papers using RAG (Retrieval-Augmented Generation)
django langchain llama2 llm pytorch reactjs transformers
Last synced: 20 Aug 2026
https://github.com/cambridgeltl/sail-bli
Self-Augmented In-Context Learning for Unsupervised Word Translation (ACL 2024). Keywords: Bilingual Lexicon Induction, Word Translation, Large Language Models, LLMs.
bilingual-dictionary-induction bilingual-lexicon-extraction bilingual-lexicon-induction few-shot-learning in-context-learning large-language-models llama llama2 llms low-resource-machine-translation machine-translation multilingual-models multilingual-nlp prompt prompt-engineering prompting pytorch self-learning word-translation zero-shot-learning
Last synced: 20 Aug 2026
https://github.com/v3xlrm1nOwo1/know-Medical-Dialogues-With-LLama2
llama2 medical python pytorch qlora transformers
Last synced: 20 Aug 2026
https://github.com/BobaZooba/shurale
Conversation AI model for open domain dialogs
bitsandbytes cerebras chatgpt deep-learning deep-neural-networks gpt gpt-4 gptq large-language-models llama llama2 llm mistral natural-language-processing nlp openai pytorch torch transformers vicuna
Last synced: 20 Aug 2026
https://github.com/FareedKhan-dev/2.3-million-parameter-llm-depriac
Building a 2.3M-parameter LLM from scratch with LLaMA 1 architecture.
deep-learning depricated huggingface large-language-models llama llm python pytorch transformers
Last synced: 20 Aug 2026
https://github.com/mohitpg/LLMs-from-scratch
A collection of LLMs implemented from scratch using pytorch
gpt llama llm llm-inference model pytorch
Last synced: 21 Aug 2026
https://github.com/Abstract-Dex/Document-Chatbot
Small document chatbot using LlamaIndex and Llama-2
chromadb langchain llama2 llamaindex pytorch sentence-transformers text-generation
Last synced: 20 Aug 2026
https://github.com/Simphoni/llama2-inference-mlu290
Distributed llama2 inference code for Cambricon MLU290. Relies only on PyTorch.
cambricon distributed-inferencing llama2 pytorch
Last synced: 20 Aug 2026
https://github.com/smendes2901/movie_chat_bot
This repository contains the implementation of a fine-tuned Llama2 chatbot using QLoRA, tailored to provide detailed information and recommendations about movies. The model is fine-tuned on the IMDB dataset, enabling it to generate informed and contextually relevant responses.
chatbot embeddings fine-tuning huggingface instruction-tuning llama2 llm pandas pytorch recommender-system sentence vector
Last synced: 20 Aug 2026
https://github.com/higgsfield-ai/higgsfield
Fault-tolerant, highly scalable GPU orchestration, and a machine learning framework designed for training models with billions to trillions of parameters
cluster-management deep-learning distributed llama llama2 llm machine-learning mlops pytorch
Last synced: 21 Aug 2026
https://github.com/piex-1/llama-pytorch
code pytorch in llama
llama llm pytorch
Last synced: 21 Aug 2026
https://github.com/AI-Hypercomputer/JetStream
JetStream is a throughput and memory optimized engine for LLM inference on XLA devices, starting with TPUs (and GPUs in future -- PRs welcome).
gemma gpt gpu inference jax large-language-models llama llama2 llm llm-inference llmops mlops model-serving pytorch tpu transformer
Last synced: 21 Aug 2026
https://github.com/prince-css/semrep
Factuality check of the SemRep Predications
bert drug-discovery drug-drug-interaction electra few-shot-learning gpt-35-turbo in-context-learning llama2 llm-finetuning longformer nlp openai peft python pytorch roberta t5
Last synced: 20 Aug 2026
https://github.com/shevtsev/WordWeaver-AI
WordWeaver-AI: A Neural Network for Text Editing
fine-tuning huggingface jupyter-notebook llama2 neural-network pytorch text-editor
Last synced: 20 Aug 2026
https://github.com/Mahesh3394/query_response_generation_llama
In this project we hosted LLAMA model with 7B parameter for response generation. Here we created a rest api which can generate a response when provided a query text.
ctransformers fastapi langchain llama2 llm nlp pytorch rest-api
Last synced: 20 Aug 2026
Statistics
- Projects: 2,239
- Last updated: about 2 years ago