LLMForge exposes an OpenAI-compatible REST API and a gRPC streaming interface.
All requests require an API key in the Authorization header.
curl https://llm.animereze.site/v1/chat/completions \ -H "Authorization: Bearer $API_KEY" \ -H "Content-Type: application/json" \ -d '{"model":"llama-3.3-70b-instruct","stream":true,"messages":[{"role":"user","content":"Hello"}]}'
| Method | Path | Description |
|---|---|---|
| POST | /v1/chat/completions | Chat completions (SSE streaming) |
| POST | /v1/embeddings | Text embeddings |
| GET | /v1/models | List available models |
| gRPC | llmforge.generate.v1.TokenStream | Bidirectional token streaming |
Limits are enforced per API key. Exceeding them returns 429 with a Retry-After header.