8d26907c8961e91314efa14f0fbc09a6400cae2e
TODO
- Race condition on jwks token refresh
- Rate Limiting
git tag -d v1.0.0; git push origin :refs/tags/v1.0.0; git tag -a v1.0.0 -m "Release v1.0.0"; git push origin v1.0.0
🦙 Ollama Rust API Wrapper
A high-performance Rust API wrapper around Ollama, providing an OpenAI-compatible interface, model lifecycle management, and advanced runtime features.
🚀 Features
- ✅ OpenAI-compatible API (
/v1/...) - ⚡ Streaming (Server-Sent Events)
- 🧠 Model lifecycle management (load/unload)
- 🔐 API key authentication (optional)
- 📊 Usage tracking & observability
- 🔀 Model routing & abstraction
- 🧩 Extensible architecture (multi-provider ready)
📡 API Endpoints
1. Core LLM API (OpenAI-compatible)
Chat Completions
POST /v1/chat/completions
Text Completions
POST /v1/completions
Embeddings
POST /v1/embeddings
List Models
GET /v1/models
2. Model Lifecycle Management
Load (Warmup)
POST /v1/models/{model}/load
{
"keep_alive": "10m"
}
Unload (Free Memory)
POST /v1/models/{model}/unload
Internally uses:
{
"model": "...",
"keep_alive": 0
}
Reload (Optional)
POST /v1/models/{model}/reload
3. Model Management
Pull Model
POST /v1/models/pull
Delete Model
DELETE /v1/models/{model}
4. Runtime & Observability
Model Status
GET /v1/models/{model}/status
List Loaded Models
GET /v1/runtime/models
5. Streaming
Enable streaming with:
{
"stream": true
}
Response format (SSE):
data: {"choices":[{"delta":{"content":"Hello"}}]}
data: {"choices":[{"delta":{"content":" world"}}]}
data: [DONE]
6. Health Checks
GET /health
GET /ready
🧠 Internal Mapping (Ollama)
| Wrapper Endpoint | Ollama Endpoint |
|---|---|
| /v1/chat/completions | /api/chat |
| /v1/completions | /api/generate |
| /v1/embeddings | /api/embeddings |
| /v1/models | /api/tags |
| /v1/models/pull | /api/pull |
| DELETE /v1/models/{model} | /api/delete |
| load/unload | /api/generate |
⚙️ Configuration
Docker (optional default)
environment:
- OLLAMA_KEEP_ALIVE=10m
Note: Request-level
keep_aliveoverrides this value.
🔧 Advanced Features
🔀 Model Routing
Use abstract model names:
{
"model": "fast"
}
Example mapping:
fast → llama3:8b
smart → llama3:70b
code → deepseek-coder
📊 Usage Tracking
GET /v1/usage
Tracks:
- request count
- latency
- per-model usage
🔐 Authentication
Authorization: Bearer sk-xxxx
Endpoints:
POST /v1/keys
DELETE /v1/keys/{id}
🚦 Rate Limiting
- Requests per minute
- Tokens per minute
Returns:
429 Too Many Requests
🧠 Sessions (Context Management)
POST /v1/sessions
POST /v1/sessions/{id}/chat
Stores conversation history server-side.
⚡ Caching
- Embeddings
- Deterministic prompts (temperature = 0)
🧩 Tool / Function Calling
Supports structured tool execution:
{
"tools": [
{
"name": "function_name",
"parameters": {}
}
]
}
📦 Batch Requests
POST /v1/batch
🧠 Auto Eviction
POST /v1/runtime/evict
Strategies:
- LRU
- memory threshold
🧾 Logs
GET /v1/logs
📚 Embedding Store (Optional)
POST /v1/documents
POST /v1/search
🔔 Async Jobs / Webhooks
POST /v1/jobs
🏗️ Architecture
Client → Rust API → Ollama → Response
Layers:
- HTTP (Axum)
- Service layer (business logic)
- Provider abstraction
- Ollama client
🔌 Provider Abstraction (Future-Proof)
trait LlmProvider {
async fn chat(...);
async fn embeddings(...);
}
Supports:
- Ollama (current)
- OpenAI (future)
- Others
⚠️ Notes
- Ollama has no native unload endpoint → simulated via
keep_alive = 0 - Streaming uses NDJSON → converted to SSE
- Chunk handling must be robust (partial JSON)
🎯 Roadmap
- Full OpenAI compatibility
- Multi-node routing
- GPU-aware scheduling
- Web UI dashboard
- Distributed inference
🧠 Summary
This project turns Ollama into:
👉 A local OpenAI-compatible API 👉 A controllable model runtime 👉 A foundation for a full LLM gateway
📜 License
MIT
Languages
Rust
99.2%
Dockerfile
0.8%