Skip to content
0x-latentPublic

About

for local llm and agent swarm

Resources

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

Local LLM API Gateway

统一的本地大模型API网关,提供OpenAI兼容接口,支持LM Studio、Ollama等本地模型服务。

功能特性

  • ✅ OpenAI兼容API接口
  • ✅ 支持流式和非流式响应
  • ✅ 自动路由到不同后端(LM Studio/Ollama/OpenAI)
  • ✅ 统一的错误处理
  • ✅ API密钥认证
  • ⏳ 模型自动发现
  • ⏳ 请求缓存
  • ⏳ 负载均衡

项目结构

LOCAL_LLM/
├── app/
│   ├── main.py              # 应用入口
│   ├── core/
│   │   ├── config.py        # 配置管理
│   │   └── llm_client.py    # LLM客户端
│   ├── api/
│   │   └── v1/
│   │       ├── chat.py      # 聊天接口
│   │       └── models.py    # 模型管理
│   └── models/
│       └── schemas.py       # 数据模型
├── requirements.txt         # 依赖包
├── .env                     # 配置文件
└── README.md

快速开始

1. 安装依赖

pip install -r requirements.txt

2. 配置环境变量

复制 .env.example 到 .env 并修改配置:

# LM Studio配置
LM_STUDIO_BASE_URL=http://localhost:1234/v1
LM_STUDIO_API_KEY=lm-studio

# API网关配置
API_PORT=8000
API_KEY=sk-local-gateway-key

3. 启动LM Studio

确保LM Studio已启动并加载了模型。

4. 启动API网关

python -m app.main

服务将在 http://localhost:8000 启动。

API使用

兼容OpenAI SDK

import openai

client = openai.OpenAI(
    base_url="http://localhost:8000/v1",
    api_key="sk-local-gateway-key"
)

# 聊天补全
response = client.chat.completions.create(
    model="qwen3-30b",
    messages=[
        {"role": "user", "content": "你好!"}
    ]
)

print(response.choices[0].message.content)

# 流式响应
stream = client.chat.completions.create(
    model="qwen3-30b",
    messages=[{"role": "user", "content": "讲个故事"}],
    stream=True
)

for chunk in stream:
    if chunk.choices[0].delta.content:
        print(chunk.choices[0].delta.content, end="")

直接HTTP请求

curl http://localhost:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer sk-local-gateway-key" \
  -d '{
    "model": "qwen3-30b",
    "messages": [{"role": "user", "content": "你好"}]
  }'

API接口

POST /v1/chat/completions

聊天补全接口(OpenAI兼容)

请求体:

{
  "model": "qwen3-30b",
  "messages": [
    {"role": "user", "content": "你好"}
  ],
  "temperature": 0.7,
  "stream": false
}

GET /v1/models

获取可用模型列表

模型路由规则

网关会根据模型名称自动选择后端:

  • qwen* → LM Studio
  • gpt-4* → OpenAI
  • ollama/* → Ollama

开发计划

  • 实现LM Studio客户端
  • 实现Ollama客户端
  • 实现模型自动发现
  • 添加请求日志
  • 添加性能监控
  • 实现请求缓存
  • 多模型并发支持

许可证

MIT

About

for local llm and agent swarm

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages