统一的本地大模型API网关,提供OpenAI兼容接口,支持LM Studio、Ollama等本地模型服务。
- ✅ OpenAI兼容API接口
- ✅ 支持流式和非流式响应
- ✅ 自动路由到不同后端(LM Studio/Ollama/OpenAI)
- ✅ 统一的错误处理
- ✅ API密钥认证
- ⏳ 模型自动发现
- ⏳ 请求缓存
- ⏳ 负载均衡
LOCAL_LLM/
├── app/
│ ├── main.py # 应用入口
│ ├── core/
│ │ ├── config.py # 配置管理
│ │ └── llm_client.py # LLM客户端
│ ├── api/
│ │ └── v1/
│ │ ├── chat.py # 聊天接口
│ │ └── models.py # 模型管理
│ └── models/
│ └── schemas.py # 数据模型
├── requirements.txt # 依赖包
├── .env # 配置文件
└── README.md
pip install -r requirements.txt复制 .env.example 到 .env 并修改配置:
# LM Studio配置
LM_STUDIO_BASE_URL=http://localhost:1234/v1
LM_STUDIO_API_KEY=lm-studio
# API网关配置
API_PORT=8000
API_KEY=sk-local-gateway-key确保LM Studio已启动并加载了模型。
python -m app.main服务将在 http://localhost:8000 启动。
import openai
client = openai.OpenAI(
base_url="http://localhost:8000/v1",
api_key="sk-local-gateway-key"
)
# 聊天补全
response = client.chat.completions.create(
model="qwen3-30b",
messages=[
{"role": "user", "content": "你好!"}
]
)
print(response.choices[0].message.content)
# 流式响应
stream = client.chat.completions.create(
model="qwen3-30b",
messages=[{"role": "user", "content": "讲个故事"}],
stream=True
)
for chunk in stream:
if chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="")curl http://localhost:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer sk-local-gateway-key" \
-d '{
"model": "qwen3-30b",
"messages": [{"role": "user", "content": "你好"}]
}'聊天补全接口(OpenAI兼容)
请求体:
{
"model": "qwen3-30b",
"messages": [
{"role": "user", "content": "你好"}
],
"temperature": 0.7,
"stream": false
}获取可用模型列表
网关会根据模型名称自动选择后端:
qwen*→ LM Studiogpt-4*→ OpenAIollama/*→ Ollama
- 实现LM Studio客户端
- 实现Ollama客户端
- 实现模型自动发现
- 添加请求日志
- 添加性能监控
- 实现请求缓存
- 多模型并发支持
MIT