Star 历史趋势
数据来源: GitHub API · 生成自 Stargazers.cn
README.md

openarc_DOOM

Discord Hugging Face Devices Ask DeepWiki Documentation

[!NOTE] OpenArc is under active development.

[!NOTE] OpenArc currently requires nightly wheels to run the latest models.

uv pip install --pre -U openvino-genai --extra-index-url https://storage.openvinotoolkit.org/simple/wheels/nightly

OpenArc is an inference engine for Intel devices.

Serve LLMs, VLMs, Whisper, Kokoro-TTS, Qwen-TTS, Qwen-ASR, Embedding and Reranker models over OpenAI compatible endpoints, powered by OpenVINO on your device. Local, private, open source AI. OpenArc enables you to host speech to text, text to speech and an LLM on the same server, at the same time.

OpenArc is a community-driven effort to make acceleration from OpenVINO easier to access, deploy and leverage for our usecases.

If you are interested in using Intel devices for AI and machine learning, feel free to stop by our Discord, where we are tracking almost the whole stack, including development of llama.cpp SYCL backend.

Thanks to everyone on Discord for their continued support!

[!NOTE] Documentation lives here

Quickstart

Features

  • Support for openvino genai scheduler_config
  • NEW! Containerization with Docker #60 by @meatposes
  • NEW! Speculative decoding support for LLMs #57 by @meatposes
  • NEW! Streaming cancellation support for LLMs and VLMs
  • Multi GPU Pipeline Paralell
  • CPU offload/Hybrid device
  • NPU device support
  • OpenAI compatible endpoints
    • /v1/models
    • /v1/completions: llm only
    • /v1/chat/completions
    • /v1/audio/transcriptions: whisper, qwen3_asr
    • /v1/audio/speech: kokoro only
    • /v1/embeddings: qwen3-embedding #33 by @mwrothbe
    • /v1/rerank: qwen3-reranker #39 by @mwrothbe
  • jinja templating with AutoTokenizers
  • OpenAI Compatible tool and reasoning parsing
  • Fully async multi engine, multi task architecture
  • Model concurrency: load and infer multiple models at once
  • Automatic unload on inference failure
  • llama-bench style benchmarking for llm w/automatic sqlite database
  • metrics on every request
    • ttft
    • prefill_throughput
    • decode_throughput
    • decode_duration
    • tpot
    • load time
    • stream mode
  • More OpenVINO examples
  • OpenVINO implementation of hexgrad/Kokoro-82M
  • OpenVINO implementation of Qwen3-TTS and Qwen3-ASR

[!NOTE] Interested in contributing? Please discuss with us on discord or open an issue before submitting a PR!

Acknowledgments

OpenArc stands on the shoulders of many other projects:

Optimum-Intel

OpenVINO

OpenVINO GenAI

llama.cpp

vLLM

Transformers

FastAPI

click

rich-click

@article{zhou2024survey,
  title={A Survey on Efficient Inference for Large Language Models},
  author={Zhou, Zixuan and Ning, Xuefei and Hong, Ke and Fu, Tianyu and Xu, Jiaming and Li, Shiyao and Lou, Yuming and Wang, Luning and Yuan, Zhihang and Li, Xiuhong and Yan, Shengen and Dai, Guohao and Zhang, Xiao-Ping and Dong, Yuhan and Wang, Yu},
  journal={arXiv preprint arXiv:2404.14294},
  year={2024}
}

Thanks for your work!!

关于 About

Inference engine for Intel devices. Serve LLMs, VLMs, Whisper, Kokoro-TTS, Embedding and Rerank models over OpenAI endpoints.
agentic-aifastapiinference-engineopenvino-genaiopenvino-toolkitoptimum-inteltransformers

语言 Languages

Python97.3%
Dockerfile1.4%
C++0.9%
Batchfile0.2%
Shell0.2%

提交活跃度 Commit Activity

代码提交热力图
过去 52 周的开发活跃度
476
Total Commits
峰值: 42次/周
Less
More

核心贡献者 Contributors