# TODO - Race condition on jwks token refresh - Rate Limiting git tag -d v1.0.0; git push origin :refs/tags/v1.0.0; git tag -a v1.0.0 -m "Release v1.0.0"; git push origin v1.0.0 # ๐Ÿฆ™ Ollama Rust API Wrapper A high-performance Rust API wrapper around Ollama, providing an OpenAI-compatible interface, model lifecycle management, and advanced runtime features. --- # ๐Ÿš€ Features * โœ… OpenAI-compatible API (`/v1/...`) * โšก Streaming (Server-Sent Events) * ๐Ÿง  Model lifecycle management (load/unload) * ๐Ÿ” API key authentication (optional) * ๐Ÿ“Š Usage tracking & observability * ๐Ÿ”€ Model routing & abstraction * ๐Ÿงฉ Extensible architecture (multi-provider ready) --- # ๐Ÿ“ก API Endpoints ## 1. Core LLM API (OpenAI-compatible) ### Chat Completions ``` POST /v1/chat/completions ``` ### Text Completions ``` POST /v1/completions ``` ### Embeddings ``` POST /v1/embeddings ``` ### List Models ``` GET /v1/models ``` --- ## 2. Model Lifecycle Management ### Load (Warmup) ``` POST /v1/models/{model}/load ``` ```json { "keep_alive": "10m" } ``` --- ### Unload (Free Memory) ``` POST /v1/models/{model}/unload ``` Internally uses: ```json { "model": "...", "keep_alive": 0 } ``` --- ### Reload (Optional) ``` POST /v1/models/{model}/reload ``` --- ## 3. Model Management ### Pull Model ``` POST /v1/models/pull ``` ### Delete Model ``` DELETE /v1/models/{model} ``` --- ## 4. Runtime & Observability ### Model Status ``` GET /v1/models/{model}/status ``` ### List Loaded Models ``` GET /v1/runtime/models ``` --- ## 5. Streaming Enable streaming with: ```json { "stream": true } ``` Response format (SSE): ``` data: {"choices":[{"delta":{"content":"Hello"}}]} data: {"choices":[{"delta":{"content":" world"}}]} data: [DONE] ``` --- ## 6. Health Checks ``` GET /health GET /ready ``` --- # ๐Ÿง  Internal Mapping (Ollama) | Wrapper Endpoint | Ollama Endpoint | | ------------------------- | --------------- | | /v1/chat/completions | /api/chat | | /v1/completions | /api/generate | | /v1/embeddings | /api/embeddings | | /v1/models | /api/tags | | /v1/models/pull | /api/pull | | DELETE /v1/models/{model} | /api/delete | | load/unload | /api/generate | --- # โš™๏ธ Configuration ### Docker (optional default) ```yaml environment: - OLLAMA_KEEP_ALIVE=10m ``` > Note: Request-level `keep_alive` overrides this value. --- # ๐Ÿ”ง Advanced Features ## ๐Ÿ”€ Model Routing Use abstract model names: ```json { "model": "fast" } ``` Example mapping: ``` fast โ†’ llama3:8b smart โ†’ llama3:70b code โ†’ deepseek-coder ``` --- ## ๐Ÿ“Š Usage Tracking ``` GET /v1/usage ``` Tracks: * request count * latency * per-model usage --- ## ๐Ÿ” Authentication ``` Authorization: Bearer sk-xxxx ``` Endpoints: ``` POST /v1/keys DELETE /v1/keys/{id} ``` --- ## ๐Ÿšฆ Rate Limiting * Requests per minute * Tokens per minute Returns: ``` 429 Too Many Requests ``` --- ## ๐Ÿง  Sessions (Context Management) ``` POST /v1/sessions POST /v1/sessions/{id}/chat ``` Stores conversation history server-side. --- ## โšก Caching * Embeddings * Deterministic prompts (temperature = 0) --- ## ๐Ÿงฉ Tool / Function Calling Supports structured tool execution: ```json { "tools": [ { "name": "function_name", "parameters": {} } ] } ``` --- ## ๐Ÿ“ฆ Batch Requests ``` POST /v1/batch ``` --- ## ๐Ÿง  Auto Eviction ``` POST /v1/runtime/evict ``` Strategies: * LRU * memory threshold --- ## ๐Ÿงพ Logs ``` GET /v1/logs ``` --- ## ๐Ÿ“š Embedding Store (Optional) ``` POST /v1/documents POST /v1/search ``` --- ## ๐Ÿ”” Async Jobs / Webhooks ``` POST /v1/jobs ``` --- # ๐Ÿ—๏ธ Architecture ``` Client โ†’ Rust API โ†’ Ollama โ†’ Response ``` ### Layers: * HTTP (Axum) * Service layer (business logic) * Provider abstraction * Ollama client --- # ๐Ÿ”Œ Provider Abstraction (Future-Proof) ```rust trait LlmProvider { async fn chat(...); async fn embeddings(...); } ``` Supports: * Ollama (current) * OpenAI (future) * Others --- # โš ๏ธ Notes * Ollama has no native unload endpoint โ†’ simulated via `keep_alive = 0` * Streaming uses NDJSON โ†’ converted to SSE * Chunk handling must be robust (partial JSON) --- # ๐ŸŽฏ Roadmap * [ ] Full OpenAI compatibility * [ ] Multi-node routing * [ ] GPU-aware scheduling * [ ] Web UI dashboard * [ ] Distributed inference --- # ๐Ÿง  Summary This project turns Ollama into: ๐Ÿ‘‰ A local OpenAI-compatible API ๐Ÿ‘‰ A controllable model runtime ๐Ÿ‘‰ A foundation for a full LLM gateway --- # ๐Ÿ“œ License MIT