跳到主要內容

Reame

Speed up your AI with Reame's CPU inference server

AI 20天前
分類
AI
網站
github.com
語言
英語
發布於
20天前
星標
106
複刻
9

swellweb/reame ↗ · C++ · MIT · 最近提交 5天前

Reame is a powerful CPU inference server designed to optimize the performance of large language models (LLMs) on your existing hardware. It addresses the common challenge of slow inference times by caching prompts, prefixes, and past generations to disk, ensuring that subsequent requests are significantly faster—request #100 costs a fraction of request #1. With an OpenAI-compatible API built on llama.cpp, Reame is accessible for various setups including free tiers, shared VPS, and 2-core ARM boxes. Ideal for developers, researchers, and businesses looking to leverage AI without investing in expensive hardware, Reame makes it easy to run efficient inference tasks seamlessly.

喜歡這個產品嗎?