LucasDLTG fa09e75a09
CI / Rust CI (push) Successful in 1m11s
Publish & Deploy / Build and Push to Registry (push) Successful in 17s
Publish & Deploy / Deploy via SSH (push) Successful in 18s
feat: use env var for ollama url
2026-04-09 21:17:29 +02:00
2026-04-09 21:17:29 +02:00
2026-04-09 16:46:00 +02:00
2026-04-09 15:31:55 +02:00
2026-04-09 15:31:55 +02:00
2026-04-09 16:50:14 +02:00
2026-04-09 16:46:00 +02:00
2026-04-09 16:52:20 +02:00
2026-04-09 15:50:59 +02:00
2026-04-09 15:50:59 +02:00
2026-04-09 20:51:20 +02:00
2026-04-09 15:50:59 +02:00

TODO

  • Race condition on jwks token refresh
  • Rate Limiting

git tag -d v1.0.0; git push origin :refs/tags/v1.0.0; git tag -a v1.0.0 -m "Release v1.0.0"; git push origin v1.0.0

🦙 Ollama Rust API Wrapper

A high-performance Rust API wrapper around Ollama, providing an OpenAI-compatible interface, model lifecycle management, and advanced runtime features.


🚀 Features

  • ✅ OpenAI-compatible API (/v1/...)
  • ⚡ Streaming (Server-Sent Events)
  • 🧠 Model lifecycle management (load/unload)
  • 🔐 API key authentication (optional)
  • 📊 Usage tracking & observability
  • 🔀 Model routing & abstraction
  • 🧩 Extensible architecture (multi-provider ready)

📡 API Endpoints

1. Core LLM API (OpenAI-compatible)

Chat Completions

POST /v1/chat/completions

Text Completions

POST /v1/completions

Embeddings

POST /v1/embeddings

List Models

GET /v1/models

2. Model Lifecycle Management

Load (Warmup)

POST /v1/models/{model}/load
{
  "keep_alive": "10m"
}

Unload (Free Memory)

POST /v1/models/{model}/unload

Internally uses:

{
  "model": "...",
  "keep_alive": 0
}

Reload (Optional)

POST /v1/models/{model}/reload

3. Model Management

Pull Model

POST /v1/models/pull

Delete Model

DELETE /v1/models/{model}

4. Runtime & Observability

Model Status

GET /v1/models/{model}/status

List Loaded Models

GET /v1/runtime/models

5. Streaming

Enable streaming with:

{
  "stream": true
}

Response format (SSE):

data: {"choices":[{"delta":{"content":"Hello"}}]}

data: {"choices":[{"delta":{"content":" world"}}]}

data: [DONE]

6. Health Checks

GET /health
GET /ready

🧠 Internal Mapping (Ollama)

Wrapper Endpoint Ollama Endpoint
/v1/chat/completions /api/chat
/v1/completions /api/generate
/v1/embeddings /api/embeddings
/v1/models /api/tags
/v1/models/pull /api/pull
DELETE /v1/models/{model} /api/delete
load/unload /api/generate

⚙️ Configuration

Docker (optional default)

environment:
  - OLLAMA_KEEP_ALIVE=10m

Note: Request-level keep_alive overrides this value.


🔧 Advanced Features

🔀 Model Routing

Use abstract model names:

{
  "model": "fast"
}

Example mapping:

fast   → llama3:8b
smart  → llama3:70b
code   → deepseek-coder

📊 Usage Tracking

GET /v1/usage

Tracks:

  • request count
  • latency
  • per-model usage

🔐 Authentication

Authorization: Bearer sk-xxxx

Endpoints:

POST /v1/keys
DELETE /v1/keys/{id}

🚦 Rate Limiting

  • Requests per minute
  • Tokens per minute

Returns:

429 Too Many Requests

🧠 Sessions (Context Management)

POST /v1/sessions
POST /v1/sessions/{id}/chat

Stores conversation history server-side.


⚡ Caching

  • Embeddings
  • Deterministic prompts (temperature = 0)

🧩 Tool / Function Calling

Supports structured tool execution:

{
  "tools": [
    {
      "name": "function_name",
      "parameters": {}
    }
  ]
}

📦 Batch Requests

POST /v1/batch

🧠 Auto Eviction

POST /v1/runtime/evict

Strategies:

  • LRU
  • memory threshold

🧾 Logs

GET /v1/logs

📚 Embedding Store (Optional)

POST /v1/documents
POST /v1/search

🔔 Async Jobs / Webhooks

POST /v1/jobs

🏗️ Architecture

Client → Rust API → Ollama → Response

Layers:

  • HTTP (Axum)
  • Service layer (business logic)
  • Provider abstraction
  • Ollama client

🔌 Provider Abstraction (Future-Proof)

trait LlmProvider {
    async fn chat(...);
    async fn embeddings(...);
}

Supports:

  • Ollama (current)
  • OpenAI (future)
  • Others

⚠️ Notes

  • Ollama has no native unload endpoint → simulated via keep_alive = 0
  • Streaming uses NDJSON → converted to SSE
  • Chunk handling must be robust (partial JSON)

🎯 Roadmap

  • Full OpenAI compatibility
  • Multi-node routing
  • GPU-aware scheduling
  • Web UI dashboard
  • Distributed inference

🧠 Summary

This project turns Ollama into:

👉 A local OpenAI-compatible API 👉 A controllable model runtime 👉 A foundation for a full LLM gateway


📜 License

MIT

S
Description
No description provided
api
Readme Apache-2.0
677 KiB
Languages
Rust 99.2%
Dockerfile 0.8%