代码库
A 2.78-trillion-parameter Kimi K3 running inference on a single CPU in 8.24 GB of RAM. Portable C99: no BLAS, no framework, no GPU.
C
avx2c99cpu-inferencedeep-learningfrom-scratchinference-enginekimi-k3linear-attentionllmllm-inferencemachine-learningmemory-efficientmixture-of-expertsmoemxfp4quantizationsimdsystems-programmingtransformerzero-dependencies
A straightforward method to reduce your LLM inference API costs and token usage.
Jupyter Notebook
aiapiartificial-intelligencegeminilarge-language-modelsllmopenai
Self improving agentic rag pipeline
Jupyter Notebook
agentic-aiaiartificial-intelligencelangchainrag
A straightforward method for training your LLM, from downloading data to generating text.
Python
geminilarge-language-modelsllmopenaitrainingtransformers
35 production-grade agentic AI architectures (Reflexion, LATS, GraphRAG, MemGPT, Voyager, BrowserAgent, ...) — a Python library and runnable textbook with multi-provider LLM support and a 17-task benchmark leaderboard.
Jupyter Notebook
agentic-aiai-agentslangchainlanggraphlangsmithllm