Collections: awesome-llama
https://github.com/sinanazem/llm-apps
This repository showcases projects leveraging advanced language models.
chatbots llama2 ollama rag
Last synced: 12 Aug 2026
https://github.com/mbzuai-oryx/VideoGPT-plus
Official Repository of paper VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding
chatbot clip dual-encoder gpt4 gpt4o image-encoder llama3 llava multimodal phi-3-mini vicuna video-chatbot video-conversation video-encoder vision-language vision-language-pretraining
Last synced: 12 Aug 2026
https://github.com/mudler/LocalAI
LocalAI is the open-source AI engine. Run any model - LLMs, vision, voice, image, video - on any hardware. No GPU required.
agents ai api audio-generation decentralized distributed image-generation libp2p llama llm mamba mcp musicgen object-detection rerank stable-diffusion text-generation tts
Last synced: 13 Aug 2026
https://github.com/PaddlePaddle/PaddleNLP
Easy-to-use and powerful LLM and SLM library with awesome model zoo.
bert compression distributed-training document-intelligence embedding ernie information-extraction llama llm neural-search nlp paddlenlp pretrained-models question-answering search-engine semantic-analysis sentiment-analysis transformers uie
Last synced: 13 Aug 2026
https://github.com/tenstorrent/tt-metal
:metal: TT-NN operator library, and TT-Metalium low level kernel programming model.
accelerator ai cuda deepseek gpu img-gen kernels llama llm metal scale-out stable-diffusion tenstorrent video-gen
Last synced: 13 Aug 2026
https://github.com/itsyaasir/pdf-intellect
PDF Intellect is a smart AI-powered tool designed to extract and analyze information from PDF document
ai analyzer conda llama2 llamacpp machine-learning pdf project python smart
Last synced: 12 Aug 2026
https://github.com/Shaun-le/ViQAG
Question and answer generation (QAG) is a natural language processing (NLP) task that generates a question and an answer in the same time by using context information. The input context can be represented in form of structured information in a database or raw text. The outputs of QAG systems can be directly applied to several NLP applications...
large-language-models llama2 pre-trained-language-models question-and-answer-generation
Last synced: 12 Aug 2026
https://github.com/marcopoli/LLaMAntino-3-ANITA
The 🌟ANITA project🌟 *(Advanced Natural-based interaction for the ITAlian language)* wants to provide Italian NLP researchers with an improved model the for Italian Language 🇮🇹 use cases.
llama3 llama3-anita llama3-finetune llama3-meta-ai llamantino llms meta-llama
Last synced: 12 Aug 2026
https://github.com/FoundationVision/Groma
[ECCV2024] Grounded Multimodal Large Language Model with Localized Visual Tokenization
foundation-models grounding large-language-models llama llama2 llm mllm multimodal vision-language-model
Last synced: 12 Aug 2026
https://github.com/zetavg/llama-lora-tuner
UI tool for fine-tuning and testing your own LoRA models base on LLaMA, GPT-J and more. One-click run on Google Colab. + A Gradio ChatGPT-like Chat UI to demonstrate your language models.
ai alpaca alpaca-lora google-colab gpt gpt-j language-model llama lora machine-learning peft
Last synced: 12 Aug 2026
https://github.com/sooraj12/ragchatbot_llama3
rag
agent langchain llama3 rag
Last synced: 12 Aug 2026
https://github.com/luogen1996/LLaVA-HR
[ICLR2025] LLaVA-HR: High-Resolution Large Language-Vision Assistant
high-resolution-mllms llama2 llms multimodal-chatbot
Last synced: 12 Aug 2026
https://github.com/MajidRaimi/Chat-With-PDF
Small project where you can chat with pdf using langchain framework.
chroma langchain llama2 llm python
Last synced: 12 Aug 2026
https://github.com/dylanhogg/llmgraph
Create knowledge graphs with LLMs
chatgpt gephi gexf graph graphml knowledge-graph large-language-model llama2 llm
Last synced: 12 Aug 2026
https://github.com/declare-lab/red-instruct
Codes and datasets of the paper Red-Teaming Large Language Models using Chain of Utterances for Safety-Alignment
huggingface-transformers llama llama2 llm llms
Last synced: 12 Aug 2026
https://github.com/c0sogi/llama-api
An OpenAI-like LLaMA inference API
api exllama fastapi llama llamacpp
Last synced: 12 Aug 2026
https://github.com/chatchat-space/langchain-chatchat
Langchain-Chatchat(原Langchain-ChatGLM)基于 Langchain 与 ChatGLM, Qwen 与 Llama 等语言模型的 RAG 与 Agent 应用 | Langchain-Chatchat (formerly langchain-ChatGLM), local knowledge based LLM (like ChatGLM, Qwen and Llama) RAG and Agent app with langchain
chatbot chatchat chatglm chatgpt embedding faiss fastchat gpt knowledge-base langchain langchain-chatglm llama llm milvus ollama qwen rag retrieval-augmented-generation streamlit xinference
Last synced: 12 Aug 2026
https://github.com/tmc/go-llama2
Llama 2 inference in one file of pure Go
golang llama2
Last synced: 12 Aug 2026
https://github.com/ATH-MaaS/Ovis
A novel Multimodal Large Language Model (MLLM) architecture, designed to structurally align visual and textual embeddings.
chatbot llama3 multimodal multimodal-large-language-models multimodality qwen vision-language-learning vision-language-model
Last synced: 12 Aug 2026
https://github.com/Bklieger/infinite-bookshelf
Infinite Bookshelf: Generate entire books in seconds using Groq and Llama3
ai groq groq-api llama3
Last synced: 12 Aug 2026
https://github.com/mudler/localai
LocalAI is the open-source AI engine. Run any model - LLMs, vision, voice, image, video - on any hardware. No GPU required.
agents ai api audio-generation decentralized distributed image-generation libp2p llama llm mamba mcp musicgen object-detection rerank stable-diffusion text-generation tts
Last synced: 13 Aug 2026
https://github.com/jindalAnuj/LLama3-AIEmailAgent
LLama3 agent using crewai , detect spam email and auto responsd if important
agent crewai llama3 python
Last synced: 12 Aug 2026
https://github.com/wxjiao/ParroT
The ParroT framework to enhance and regulate the Translation Abilities during Chat based on open-sourced LLMs (e.g., LLaMA-7b, Bloomz-7b1-mt) and human written translation and evaluation data.
bloomz chatgpt contrastive error-guided gpt-4 human-feedback instruction-tuning llama lora machine-translation
Last synced: 12 Aug 2026
https://github.com/HROlive/Poland-End-To-End-LLM-Bootcamp
This bootcamp is designed to give NLP researchers an end-to-end overview on the fundamentals of NVIDIA NeMo framework, complete solution for building large language models. It will also have hands-on exercises complimented by tutorials, code snippets, and presentations to help researchers kick-start with NeMo LLM Service and Guardrails.
gpt llama2 llm llm-inference llm-training nemo-guardrails nvidia nvidia-nemo p-tuning prompt-tuning tensorrt triton
Last synced: 12 Aug 2026
https://github.com/IAAR-Shanghai/Grimoire
Grimoire is All You Need for Enhancing Large Language Models
baichuan chatgpt datasets gpt-4 grimoire icl in-context-learning llama llm phi2
Last synced: 12 Aug 2026
https://github.com/jianzhnie/Open-R1
The open source implementation of DeepSeek-R1. 开源复现 DeepSeek-R1
deepseek-r1 deepseek-v3 grpo llm rlhf
Last synced: 12 Aug 2026
https://github.com/X-PLUG/mPLUG-Owl
mPLUG-Owl: The Powerful Multi-modal Large Language Model Family
alpaca chatbot chatgpt damo dialogue gpt gpt4 gpt4-api huggingface instruction-tuning large-language-models llama mplug mplug-owl multimodal pretraining pytorch transformer video visual-recognition
Last synced: 12 Aug 2026
https://github.com/mddunlap924/PII-Detection
Personal Identifiable Information (PII) entity detection and performance enhancement with synthetic data generation
fine-tuning kaggle-competition llama3 llm named-entity-recognition ner pii pii-data pii-detection prompt-engineering
Last synced: 12 Aug 2026
https://github.com/farhan0167/llama-engine
High level Python API to run open source LLM models on Colab with less code
llama2 llm
Last synced: 12 Aug 2026
https://github.com/felipelmc/SciELO-Summarizer
[EN-US] Scrapes papers from SciELO and summarizes them using Llama3 | [PT-BR] Extrai artigos do SciELO e os resume utilizando Llama3
computational-social-sciences llama3 llm papers science-research webscraping
Last synced: 12 Aug 2026
https://github.com/tien02/llm-math
Fine tune Large Language Model on Mathematic dataset
huggingface llama llama2 llm lora mathematics supervised-finetuning transformer
Last synced: 12 Aug 2026
https://github.com/znx-x/arcadia-llm-server
This repository is an LLM interface based on llama.cpp and allows you to run your own Arcadia LLM server locally in just a few easy-to-follow steps.
ai bot chat cpp free gpt llama local mistral open-source server
Last synced: 12 Aug 2026
https://github.com/Tongjilibo/bert4torch
An elegent pytorch implement of transformers
belle bert bert4keras bert4torch chatglm large-language-models llama llm named-entity-recognition nlp pytorch relation-extraction seq2seq text-classification transformers
Last synced: 12 Aug 2026
https://github.com/tongjilibo/bert4torch
An elegent pytorch implement of transformers
belle bert bert4keras bert4torch chatglm large-language-models llama llm named-entity-recognition nlp pytorch relation-extraction seq2seq text-classification transformers
Last synced: 12 Aug 2026
https://github.com/rbourgeat/ImpAI
😈 ImpAI is an advanced role play app using large language and diffusion models.
ai character-ai chat docker game ggml gguf linux llama llama-cpp llm macos roleplay stable-diffusion windows
Last synced: 12 Aug 2026
https://github.com/Yxxxb/VoCo-LLaMA
[CVPR'2025] VoCo-LLaMA: This repo is the official implementation of "VoCo-LLaMA: Towards Vision Compression with Large Language Models".
image-compression llama llava
Last synced: 12 Aug 2026
https://github.com/nandxorandor/chatbot_VS_chatbot
This project showcases engaging interactions between two AI chatbots.
ai-automation chat chatbot chatbot-development chatgpt conversational-ai flask llama model python-chatbot zephyr
Last synced: 12 Aug 2026
https://github.com/AkiRusProd/llm-agent
LLM using long-term memory through vector database
chromadb gpt gpt4all intent-classification large-language-models llama llm machine-learning nlp rag
Last synced: 12 Aug 2026
https://github.com/LennardZuendorf/thesis-webapp
Webapp/Application implemention of my thesis about XAI and Interpretability of Transformer Models.
bertviz gradio huggingface interpretable-ai llama2 mistral shap xai
Last synced: 12 Aug 2026
https://github.com/riccardomusmeci/mlx-llm
Large Language Models (LLMs) applications and tools running on Apple Silicon in real-time with Apple MLX.
llama llm mistral mlx phi transformers
Last synced: 12 Aug 2026
https://github.com/MOONLAPSED/cognosis
Morphological Source Code (+QSD, /MOONLAPSED/cognosis branch) implemented in Python3 for contemporary hardware. Operates as a quantized kernel of agentic motility, akin to a Hilbert space kernel; augmented by an AdS/CFT Noetherian jet space enabling category-theoretic syntax-lift/lower, morphological differentiation, and morphosemantic integration.
cohomology compiler-theory holography
Last synced: 12 Aug 2026
https://github.com/CMKRG/QiZhenGPT
QiZhenGPT: An Open Source Chinese Medical Large Language Model|一个开源的中文医疗大语言模型
alpaca chatgpt chinese-medical disease gpt llama llm medicine
Last synced: 12 Aug 2026
https://github.com/YY0649/ICE-PIXIU
ICE-PIXIU:A Cross-Language Financial Megamodeling Framework
internlm large-language-models llama nlp pixiu vllm
Last synced: 12 Aug 2026
https://github.com/aakarsh/rl-llm-calibration-test
Attempt at replication of the parts of the paper "Language models (mostly) know what they know", on open datasets, and models.
llama2 llm
Last synced: 12 Aug 2026
https://github.com/AdrianBZG/llama-multimodal-vqa
Multimodal Instruction Tuning for Llama 3
chatbot chatgpt gpt-4 huggingface instruction-tuning language-models llama llama2 llama3 multimodal multimodal-instruction-tuning visual-language-learning visual-question-answering vqa
Last synced: 12 Aug 2026
https://github.com/jy-yuan/KIVI
[ICML 2024] KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache
inference large-language-models llama llm natural-language-processing quantization transformer
Last synced: 12 Aug 2026
https://github.com/OpenGVLab/InternGPT
InternGPT (iGPT) is an open source demo platform where you can easily showcase your AI models. Now it supports DragGAN, ChatGPT, ImageBind, multimodal chat like GPT-4, SAM, interactive image editing, etc. Try it at igpt.opengvlab.com (支持DragGAN、ChatGPT、ImageBind、SAM的在线Demo系统)
chatgpt click draggan foundation-model gpt gpt-4 gradio husky image-captioning imagebind internimage langchain llama llm multimodal sam segment-anything vicuna video-generation vqa
Last synced: 12 Aug 2026
https://github.com/anslin-raj/anil_ai_chat_llama_2
Anil AI is an innovative chat application that harnesses the robust capabilities of Generative AI, utilizing the advanced LLAMA 2 model for inference.
automation chatbot devops docker generative-ai gpt llama2 llm ml mlops
Last synced: 12 Aug 2026
https://github.com/ModelTC/lightllm
LightLLM is a Python-based LLM (Large Language Model) inference and serving framework, notable for its lightweight design, easy scalability, and high-speed performance.
deep-learning gpt llama llm model-serving nlp openai-triton
Last synced: 12 Aug 2026
https://github.com/ASSERT-KTH/repairllama
RepairLLaMA: Efficient Representations and Fine-Tuned Adapters for Program Repair http://arxiv.org/pdf/2312.15698
apr codellama llama llms lora repair
Last synced: 12 Aug 2026
https://github.com/bolna-ai/bolna
Conversational voice AI agents
agentic-ai agents ai-agents cartesia conversational-ai deepgram deepseek deepseek-chat elevenlabs function-calling gpt-4 llama openai plivo twilio voice-agents voice-ai-agents voice-assistant whisper
Last synced: 12 Aug 2026
https://github.com/thejasmeetsingh/moody-llm
A LLM whose mood keeps changing.
asynchronous-programming fastapi genai highlightjs langchain llama3 llm ollama pydantic python3 reactjs restful-api supabase tailwindcss websockets
Last synced: 12 Aug 2026
https://github.com/ROIM1998/APT
[ICML'24 Oral] APT: Adaptive Pruning and Tuning Pretrained Language Models for Efficient Training and Inference
bert efficient-deep-learning llama2 llm llm-finetuning peft peft-fine-tuning-llm pruning roberta t5
Last synced: 12 Aug 2026
https://github.com/alibaba/rtp-llm
RTP-LLM: Alibaba's high-performance LLM inference engine for diverse applications.
gpt inference llama llm llm-serving llmops model-serving
Last synced: 12 Aug 2026
https://github.com/IAAR-Shanghai/FastMem
Fast Memorization of Prompt Improves Context Awareness of Large Language Models (Findings of EMNLP 2024)
context-awareness llama3 llm qwen1-5
Last synced: 12 Aug 2026
https://github.com/KompleteAI/xllm
🦖 X—LLM: Simple & Cutting Edge LLM Finetuning
alpaca cerebras chatgpt deep-learning deep-neural-networks deeplearning falcon gpt language-model large-language-models llama2 llama2-7b llm mistal mistral mistralai natural-language-processing nlp openai vicuna
Last synced: 12 Aug 2026
https://github.com/gotzmann/booster
Booster - open accelerator for LLM models. Better inference and debugging for AI hackers
chatgpt exllama ggml gpt llama llama-cpp llamacpp llm ollama oobabooga openai vllm
Last synced: 12 Aug 2026
https://github.com/allenai/natural-instructions
Expanding natural instructions
Last synced: 12 Aug 2026
https://github.com/silvanmelchior/IncognitoPilot
An AI code interpreter for sensitive data, powered by GPT-4 or Code Llama / Llama 2.
ai chatgpt chatgpt-code-interpreter codellama copilot gpt-4 llama2 llm python
Last synced: 12 Aug 2026
https://github.com/karpathy/llama2.c
Inference Llama 2 in one file of pure C
Last synced: 12 Aug 2026
https://github.com/zetavg/LLaMA-LoRA-Tuner
UI tool for fine-tuning and testing your own LoRA models base on LLaMA, GPT-J and more. One-click run on Google Colab. + A Gradio ChatGPT-like Chat UI to demonstrate your language models.
ai alpaca alpaca-lora google-colab gpt gpt-j language-model llama lora machine-learning peft
Last synced: 12 Aug 2026
https://github.com/flojud/DocsChat
The chatbot utilizes a conversational retrieval chain to answer user queries based on the content of embedded documents. It leverages various NLP techniques, including language models and embeddings, to provide relevant responses.
Last synced: 12 Aug 2026
https://github.com/time-series-foundation-models/lag-llama
Lag-Llama: Towards Foundation Models for Probabilistic Time Series Forecasting
forecasting foundation-models lag-llama llama time-series time-series-forecasting time-series-prediction time-series-transformer timeseries timeseries-forecasting transformers
Last synced: 12 Aug 2026
https://github.com/Tiiny-AI/PowerInfer
High-speed Large Language Model Serving for Local Deployment
large-language-models llama llm llm-inference local-inference
Last synced: 12 Aug 2026
https://github.com/ddh0/easy-llama
Python package wrapping llama.cpp for on-device LLM inference
llama llamacpp llm llms text-generation
Last synced: 12 Aug 2026
https://github.com/eidolon-ai/eidolon
The first AI Agent Server, Eidolon is a pluggable Agent SDK and enterprise ready, deployment server for Agentic applications
agents generative-ai langchain llama llm openai python services
Last synced: 12 Aug 2026
https://github.com/xorbitsai/inference
Swap GPT for any LLM by changing a single line of code. Xinference lets you run open-source, speech, and multimodal models on cloud, on-prem, or your laptop — all through one unified, production-ready inference API.
artificial-intelligence chatglm deployment flan-t5 gemma ggml glm4 inference llama llama3 llamacpp llm machine-learning mistral openai-api pytorch qwen vllm whisper wizardlm
Last synced: 13 Aug 2026
https://github.com/horenbergerb/llamagotchi
A bunch of LLaMa model investigations, including recreating generative agents (from the paper Generative Agents: Interactive Simulacra of Human Behavior)
generative-agents langchain llama llamacpp llm machine-learning
Last synced: 12 Aug 2026
https://github.com/michaelnny/RAG-LLaMA
A clean and simple implementation of Retrieval Augmented Generation (RAG) to enhanced LLaMA chat model to answer questions from a private knowledge base. We use Tesla user manuals to build the knowledge base, and use open-source embedding and Cross-Encoders reranking models from Sentence Transformers in this project.
llama2 llm rag
Last synced: 12 Aug 2026
https://github.com/alta3/llm-the-alta3-way
The greatest LLMs on the planet!
gpt llama llm
Last synced: 12 Aug 2026
https://github.com/fishaudio/fish-speech
SOTA Open Source TTS
llama transformer tts valle vits vqgan vqvae
Last synced: 12 Aug 2026
https://github.com/janhq/jan
Jan is an open source alternative to ChatGPT that runs 100% offline on your computer.
chatgpt gpt llamacpp llm localai open-source self-hosted tauri
Last synced: 12 Aug 2026
https://github.com/X-rayLaser/DistributedLLM
Run LLM inference by spliting models into parts and hosting each part on a separate machine. Project is no longer maintained.
artificial-intelligence deep-learning large-language-models llama llamacpp machine-learning neural-networks nlp
Last synced: 12 Aug 2026
https://github.com/artitw/text2text
Text2Text Language Modeling Toolkit
chatbot chatgpt cross-lingual embeddings information-retrieval levenshtein-distance llama llm multi-lingual nlp question-generation rag search tf-idf tokenizer transformers translator
Last synced: 12 Aug 2026
https://github.com/albertstarfield/project-zephyrine
Zephyrine: An augmented Agentic Assistant GNC framework system for experimental Aircraft. Engineered for processing of navigation loops and sensor fusion and adaptation to bridge the gap between strategic virtual path and action planning and real-world kinematic control.
adacore assistant aviation-tools control-systems flightsystem ifcs openai-api realtime uav
Last synced: 12 Aug 2026
https://github.com/Azure-Samples/miyagi
Sample to envision intelligent apps with Microsoft's Copilot stack for AI-infused product experiences.
agents aks assistants azure azure-openai azureai copilot gpt-4 guidance langchain llama-index llama2 openai phi-2 prompt-engineering promptflow semantic-kernel taskweaver typechat
Last synced: 12 Aug 2026
https://github.com/dropbox/hqq
Official implementation of Half-Quadratic Quantization (HQQ)
llm machine-learning quantization
Last synced: 12 Aug 2026
https://github.com/pjlab-sys4nlp/llama-moe
⛷️ LLaMA-MoE: Building Mixture-of-Experts from LLaMA with Continual Pre-training (EMNLP 2024)
continual-pre-training expert-partition llama llm mixture-of-experts moe
Last synced: 12 Aug 2026
https://github.com/b4rtaz/distributed-llama
Distributed LLM inference. Connect home devices into a powerful cluster to accelerate LLM inference. More devices means faster inference.
distributed-computing distributed-llm llama2 llama3 llm llm-inference llms neural-network open-llm
Last synced: 12 Aug 2026
https://github.com/huggingface/text-generation-inference
Large Language Model Text Generation Inference
bloom deep-learning falcon gpt inference nlp pytorch starcoder transformer
Last synced: 12 Aug 2026
https://github.com/meta-llama/llama-cookbook
Welcome to the Llama Cookbook! This is your go to guide for Building with Llama: Getting started with Inference, Fine-Tuning, RAG. We also show you how to solve end to end problems using Llama model family and using them on various provider services
ai finetuning langchain llama llama2 llm machine-learning python pytorch vllm
Last synced: 12 Aug 2026
https://github.com/AI-Hypercomputer/JetStream
JetStream is a throughput and memory optimized engine for LLM inference on XLA devices, starting with TPUs (and GPUs in future -- PRs welcome).
gemma gpt gpu inference jax large-language-models llama llama2 llm llm-inference llmops mlops model-serving pytorch tpu transformer
Last synced: 12 Aug 2026
https://github.com/chatchat-space/Langchain-Chatchat
Langchain-Chatchat(原Langchain-ChatGLM)基于 Langchain 与 ChatGLM, Qwen 与 Llama 等语言模型的 RAG 与 Agent 应用 | Langchain-Chatchat (formerly langchain-ChatGLM), local knowledge based LLM (like ChatGLM, Qwen and Llama) RAG and Agent app with langchain
chatbot chatchat chatglm chatgpt embedding faiss fastchat gpt knowledge-base langchain langchain-chatglm llama llm milvus ollama qwen rag retrieval-augmented-generation streamlit xinference
Last synced: 12 Aug 2026
https://github.com/lenML/Speech-AI-Forge
🍦 Speech-AI-Forge is a project developed around TTS generation model, implementing an API Server and a Gradio-based WebUI.
agent asr chattts chattts-forge chinese colab cosy-voice cosyvoice english firered fireredtts fish-speech gpt llama llm ssml stt text-to-speech tts whisper
Last synced: 12 Aug 2026
https://github.com/zjunlp/EasyEdit
[ACL 2024] An Easy-to-use Knowledge Editing Framework for LLMs.
artificial-intelligence baichuan chatgpt cknowedit easyedit easyedit2 efficient gpt knowedit knowledge-editing knowlm large-language-models llama mmedit model-editing natural-language-processing safeedit tool trustworthy-ai unlearning
Last synced: 12 Aug 2026
https://github.com/yas-sim/openvino_japanese_chatbot_youri-7b-chat
WebUI LLM Japanese chatbot demo using Intel OpenVINO toolkit. This program uses rinna/youri-7b-chat model developed by Rinna Co.,Ltd..
chat chat-application chatbot chatgpt deep-learning gpt japanese large-language-model llama2 llm natural-language natural-language-processing openvino rinna youri
Last synced: 12 Aug 2026
https://github.com/thomas-yanxin/KarmaVLM
🧘🏻♂️KarmaVLM (相生):A family of high efficiency and powerful visual language model.
llama2 llava multimodel qwen2 vision-language-model visual-language-learning vlm
Last synced: 12 Aug 2026
https://github.com/icip-cas/ChatAlpaca
A Multi-Turn Dialogue Corpus based on Alpaca Instructions
Last synced: 12 Aug 2026
https://github.com/Agora-Lab-AI/Atom
a suite of finetuned LLMs for atomically precise function calling 🧪
ai artificial-intelligence convolutional-neural-networks function-calling gpt-4 llama llama2 llamacpp ml multi-modal open-source rpa rpc task-automation tool-usage transformer workflow-automation
Last synced: 12 Aug 2026
https://github.com/vllm-project/vllm
A high-throughput and memory-efficient inference and serving engine for LLMs
amd blackwell cuda deepseek deepseek-v3 gpt gpt-oss inference kimi llama llm llm-serving model-serving moe openai pytorch qwen qwen3 tpu transformer
Last synced: 13 Aug 2026
https://github.com/dhdbsrlw/Instruct-Tune-LLaMA-with-PEFT-Techniques
COSE474 DL Project (Prof. Hyun-Woo Kim)
attention-mechanism lit-llama llama llama-adapter llm peft xformers
Last synced: 12 Aug 2026
https://github.com/abhiverse01/TinyLlama-Base-CodeGen
This repository contains code and resources for fine-tuning a Tiny Llama base model for code generation tasks. The project aims to leverage advanced machine learning techniques to improve the performance of code generation models.
jupyter-notebook llama2 python3
Last synced: 12 Aug 2026
https://github.com/DeveshParagiri/sage
Personalized AI Chatbot based on Llama-2-7B-Chat model. Finetuned with personal data and quantized for optimized CPU runtime.
faiss llama2 llamacpp llm-chain llms quantization
Last synced: 12 Aug 2026
https://github.com/calcuis/llama-core
solo connector core built on llama.cpp
core gguf llama
Last synced: 12 Aug 2026
https://github.com/EvilFreelancer/benchmarking-llms
Comprehensive benchmarks and evaluations of Large Language Models (LLMs) with a focus on hardware usage, generation speed, and memory requirements.
benchmark llama llm mgpt mpt rugpt
Last synced: 12 Aug 2026
https://github.com/akanyaani/miniLLAMA
A simplified LLAMA implementation for training and inference tasks.
llama llama-models llama2 llm minillama pytorch simple-llama
Last synced: 12 Aug 2026
https://github.com/EmbeddedLLM/embeddedllm
EmbeddedLLM: API server for Embedded Device Deployment. Currently support CUDA/OpenVINO/IpexLLM/DirectML/CPU
aipc cpu directml directx-12 gemma ipexllm llama llm llm-inference llm-serving mistral model-inference npu open-source-llm openvino openvino-inference-engine phi-3 windows
Last synced: 12 Aug 2026
https://github.com/microsoft/sarathi-serve
A low-latency & high-throughput serving engine for LLMs
llama llm-inference pytorch transformer
Last synced: 12 Aug 2026
https://github.com/UgurkanTech/ArchNetAI
ArchNetAI is a Python library that leverages the Ollama API for generating AI-powered content.
ai json-schema llama llava ollama ollama-api phi3 python
Last synced: 12 Aug 2026
Statistics
- Projects: 2,239
- Last updated: about 2 years ago