feat: add models list endpoint
This commit is contained in:
@@ -1,6 +1,396 @@
|
||||
# TODO
|
||||
|
||||
- Race condition on jwks token refresh
|
||||
- Rate Limiting
|
||||
|
||||
|
||||
git tag -d v1.0.0; git push origin :refs/tags/v1.0.0; git tag -a v1.0.0 -m "Release v1.0.0"; git push origin v1.0.0
|
||||
git tag -d v1.0.0; git push origin :refs/tags/v1.0.0; git tag -a v1.0.0 -m "Release v1.0.0"; git push origin v1.0.0
|
||||
|
||||
|
||||
# 🦙 Ollama Rust API Wrapper
|
||||
|
||||
A high-performance Rust API wrapper around Ollama, providing an OpenAI-compatible interface, model lifecycle management, and advanced runtime features.
|
||||
|
||||
---
|
||||
|
||||
# 🚀 Features
|
||||
|
||||
* ✅ OpenAI-compatible API (`/v1/...`)
|
||||
* ⚡ Streaming (Server-Sent Events)
|
||||
* 🧠 Model lifecycle management (load/unload)
|
||||
* 🔐 API key authentication (optional)
|
||||
* 📊 Usage tracking & observability
|
||||
* 🔀 Model routing & abstraction
|
||||
* 🧩 Extensible architecture (multi-provider ready)
|
||||
|
||||
---
|
||||
|
||||
# 📡 API Endpoints
|
||||
|
||||
## 1. Core LLM API (OpenAI-compatible)
|
||||
|
||||
### Chat Completions
|
||||
|
||||
```
|
||||
POST /v1/chat/completions
|
||||
```
|
||||
|
||||
### Text Completions
|
||||
|
||||
```
|
||||
POST /v1/completions
|
||||
```
|
||||
|
||||
### Embeddings
|
||||
|
||||
```
|
||||
POST /v1/embeddings
|
||||
```
|
||||
|
||||
### List Models
|
||||
|
||||
```
|
||||
GET /v1/models
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 2. Model Lifecycle Management
|
||||
|
||||
### Load (Warmup)
|
||||
|
||||
```
|
||||
POST /v1/models/{model}/load
|
||||
```
|
||||
|
||||
```json
|
||||
{
|
||||
"keep_alive": "10m"
|
||||
}
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
### Unload (Free Memory)
|
||||
|
||||
```
|
||||
POST /v1/models/{model}/unload
|
||||
```
|
||||
|
||||
Internally uses:
|
||||
|
||||
```json
|
||||
{
|
||||
"model": "...",
|
||||
"keep_alive": 0
|
||||
}
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
### Reload (Optional)
|
||||
|
||||
```
|
||||
POST /v1/models/{model}/reload
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 3. Model Management
|
||||
|
||||
### Pull Model
|
||||
|
||||
```
|
||||
POST /v1/models/pull
|
||||
```
|
||||
|
||||
### Delete Model
|
||||
|
||||
```
|
||||
DELETE /v1/models/{model}
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 4. Runtime & Observability
|
||||
|
||||
### Model Status
|
||||
|
||||
```
|
||||
GET /v1/models/{model}/status
|
||||
```
|
||||
|
||||
### List Loaded Models
|
||||
|
||||
```
|
||||
GET /v1/runtime/models
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 5. Streaming
|
||||
|
||||
Enable streaming with:
|
||||
|
||||
```json
|
||||
{
|
||||
"stream": true
|
||||
}
|
||||
```
|
||||
|
||||
Response format (SSE):
|
||||
|
||||
```
|
||||
data: {"choices":[{"delta":{"content":"Hello"}}]}
|
||||
|
||||
data: {"choices":[{"delta":{"content":" world"}}]}
|
||||
|
||||
data: [DONE]
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 6. Health Checks
|
||||
|
||||
```
|
||||
GET /health
|
||||
GET /ready
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
# 🧠 Internal Mapping (Ollama)
|
||||
|
||||
| Wrapper Endpoint | Ollama Endpoint |
|
||||
| ------------------------- | --------------- |
|
||||
| /v1/chat/completions | /api/chat |
|
||||
| /v1/completions | /api/generate |
|
||||
| /v1/embeddings | /api/embeddings |
|
||||
| /v1/models | /api/tags |
|
||||
| /v1/models/pull | /api/pull |
|
||||
| DELETE /v1/models/{model} | /api/delete |
|
||||
| load/unload | /api/generate |
|
||||
|
||||
---
|
||||
|
||||
# ⚙️ Configuration
|
||||
|
||||
### Docker (optional default)
|
||||
|
||||
```yaml
|
||||
environment:
|
||||
- OLLAMA_KEEP_ALIVE=10m
|
||||
```
|
||||
|
||||
> Note: Request-level `keep_alive` overrides this value.
|
||||
|
||||
---
|
||||
|
||||
# 🔧 Advanced Features
|
||||
|
||||
## 🔀 Model Routing
|
||||
|
||||
Use abstract model names:
|
||||
|
||||
```json
|
||||
{
|
||||
"model": "fast"
|
||||
}
|
||||
```
|
||||
|
||||
Example mapping:
|
||||
|
||||
```
|
||||
fast → llama3:8b
|
||||
smart → llama3:70b
|
||||
code → deepseek-coder
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 📊 Usage Tracking
|
||||
|
||||
```
|
||||
GET /v1/usage
|
||||
```
|
||||
|
||||
Tracks:
|
||||
|
||||
* request count
|
||||
* latency
|
||||
* per-model usage
|
||||
|
||||
---
|
||||
|
||||
## 🔐 Authentication
|
||||
|
||||
```
|
||||
Authorization: Bearer sk-xxxx
|
||||
```
|
||||
|
||||
Endpoints:
|
||||
|
||||
```
|
||||
POST /v1/keys
|
||||
DELETE /v1/keys/{id}
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 🚦 Rate Limiting
|
||||
|
||||
* Requests per minute
|
||||
* Tokens per minute
|
||||
|
||||
Returns:
|
||||
|
||||
```
|
||||
429 Too Many Requests
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 🧠 Sessions (Context Management)
|
||||
|
||||
```
|
||||
POST /v1/sessions
|
||||
POST /v1/sessions/{id}/chat
|
||||
```
|
||||
|
||||
Stores conversation history server-side.
|
||||
|
||||
---
|
||||
|
||||
## ⚡ Caching
|
||||
|
||||
* Embeddings
|
||||
* Deterministic prompts (temperature = 0)
|
||||
|
||||
---
|
||||
|
||||
## 🧩 Tool / Function Calling
|
||||
|
||||
Supports structured tool execution:
|
||||
|
||||
```json
|
||||
{
|
||||
"tools": [
|
||||
{
|
||||
"name": "function_name",
|
||||
"parameters": {}
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 📦 Batch Requests
|
||||
|
||||
```
|
||||
POST /v1/batch
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 🧠 Auto Eviction
|
||||
|
||||
```
|
||||
POST /v1/runtime/evict
|
||||
```
|
||||
|
||||
Strategies:
|
||||
|
||||
* LRU
|
||||
* memory threshold
|
||||
|
||||
---
|
||||
|
||||
## 🧾 Logs
|
||||
|
||||
```
|
||||
GET /v1/logs
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 📚 Embedding Store (Optional)
|
||||
|
||||
```
|
||||
POST /v1/documents
|
||||
POST /v1/search
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 🔔 Async Jobs / Webhooks
|
||||
|
||||
```
|
||||
POST /v1/jobs
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
# 🏗️ Architecture
|
||||
|
||||
```
|
||||
Client → Rust API → Ollama → Response
|
||||
```
|
||||
|
||||
### Layers:
|
||||
|
||||
* HTTP (Axum)
|
||||
* Service layer (business logic)
|
||||
* Provider abstraction
|
||||
* Ollama client
|
||||
|
||||
---
|
||||
|
||||
# 🔌 Provider Abstraction (Future-Proof)
|
||||
|
||||
```rust
|
||||
trait LlmProvider {
|
||||
async fn chat(...);
|
||||
async fn embeddings(...);
|
||||
}
|
||||
```
|
||||
|
||||
Supports:
|
||||
|
||||
* Ollama (current)
|
||||
* OpenAI (future)
|
||||
* Others
|
||||
|
||||
---
|
||||
|
||||
# ⚠️ Notes
|
||||
|
||||
* Ollama has no native unload endpoint → simulated via `keep_alive = 0`
|
||||
* Streaming uses NDJSON → converted to SSE
|
||||
* Chunk handling must be robust (partial JSON)
|
||||
|
||||
---
|
||||
|
||||
# 🎯 Roadmap
|
||||
|
||||
* [ ] Full OpenAI compatibility
|
||||
* [ ] Multi-node routing
|
||||
* [ ] GPU-aware scheduling
|
||||
* [ ] Web UI dashboard
|
||||
* [ ] Distributed inference
|
||||
|
||||
---
|
||||
|
||||
# 🧠 Summary
|
||||
|
||||
This project turns Ollama into:
|
||||
|
||||
👉 A local OpenAI-compatible API
|
||||
👉 A controllable model runtime
|
||||
👉 A foundation for a full LLM gateway
|
||||
|
||||
---
|
||||
|
||||
# 📜 License
|
||||
|
||||
MIT
|
||||
|
||||
Reference in New Issue
Block a user