API
A program can use LLeMbas as the person it acts for: their models through an OpenAI-compatible API, their library over MCP, and — for an administrator — accounts, usage and the audit log.
Tokens
Make a token under Settings → Security → API tokens: a name, an expiry (never, 30 or 90 days, a year) and its scopes. You need the Use the API permission. The value is shown once; afterwards only its first characters, to tell tokens apart. Send it as Authorization: Bearer lmb_….
| Scope | Allows | Needs |
|---|---|---|
models | /v1: your models, chat, embeddings, audio, images, usage; search and fetch | Use the API |
library | /mcp: your memory, notes, skills and knowledge | Use the API |
link | being driven from the server as a linked device | Use the API |
admin:read | accounts and usage, read only | administrator |
admin:audit | the audit log | administrator, or Read the audit log |
admin:credits | top-ups, adjustments, plans | administrator, or Grant credits |
- Scopes are checked against the holder on every request. Take Use the API away from someone and all their tokens stop at once.
- A token is never accepted on a web page, and a browser session is never accepted on the API.
- Each token may make a limited number of requests a minute (120 by default, set in Admin → General); over it the answer is
429withRetry-After. - Your devices and tokens are listed under Settings; an administrator sees everyone’s on each user’s page.
Endpoints
| Endpoint | What |
|---|---|
GET /v1/models | The models you could start a chat with, as provider/model, each with a kind: chat, embedding, stt, tts or image — plus context, output limit, efforts, vision and tools. |
POST /v1/chat/completions | OpenAI’s chat completions, streamed or not. |
POST /v1/embeddings | Embeddings from an embedding model. |
POST /v1/audio/transcriptions | Speech to text: a multipart file, optionally language. Uses the server’s transcription model. |
POST /v1/audio/speech | Text to speech: input, optionally voice, speed, response_format (mp3, wav, flac, aac). |
POST /v1/images/generations | Image generation through the server’s ComfyUI. |
GET /v1/usage | This month’s tokens and credits, your plan and balance, and this token’s share. |
POST /api/v1/search | {"query": …, "max_results": n} → results with title, url and snippet. |
POST /api/v1/fetch | {"url": …} → the page’s title and text. |
GET /api/v1/instance | The server’s name, your account, your default model, and which services you may use. |
POST /mcp | Your library over MCP (streamable HTTP), with a library token. |
The API is stateless: nothing is kept, no chat and no message. Each request goes to the model’s connection with that connection’s own key, which never leaves the server. Requests are counted and limited like replies in the interface — tokens a month, replies at once, credits — and errors come back in OpenAI’s shape. An endpoint whose feature is switched off for you answers 404.
Examples
# list the models you may use
curl -s https://your-server/v1/models -H "Authorization: Bearer $TOKEN"
# chat, streamed
curl -s https://your-server/v1/chat/completions \
-H "Authorization: Bearer $TOKEN" -H "Content-Type: application/json" \
-d '{"model": "llama/qwen", "messages": [{"role": "user", "content": "Hello"}], "stream": true}'
# embeddings
curl -s https://your-server/v1/embeddings \
-H "Authorization: Bearer $TOKEN" -H "Content-Type: application/json" \
-d '{"model": "llama/bge-m3", "input": "waybread"}'
# transcribe, then speak
curl -s https://your-server/v1/audio/transcriptions \
-H "Authorization: Bearer $TOKEN" -F file=@note.wav
curl -s https://your-server/v1/audio/speech \
-H "Authorization: Bearer $TOKEN" -H "Content-Type: application/json" \
-d '{"input": "Speak, friend, and enter."}' -o reply.mp3
# draw
curl -s https://your-server/v1/images/generations \
-H "Authorization: Bearer $TOKEN" -H "Content-Type: application/json" \
-d '{"prompt": "a mallorn leaf, pixel art"}'
# this month's usage
curl -s https://your-server/v1/usage -H "Authorization: Bearer $TOKEN"Any OpenAI client library works: set its base URL to https://your-server/v1 and its API key to your token.
from openai import OpenAI
client = OpenAI(base_url="https://your-server/v1", api_key="lmb_…")
reply = client.chat.completions.create(
model="llama/qwen",
messages=[{"role": "user", "content": "Hello"}],
)
print(reply.choices[0].message.content)Your library over MCP
/mcp serves the same library tools a chat has — memory, notes, skills, knowledge — as far as your permissions allow, to any MCP client with a library token. One JSON-RPC message per POST, answered with JSON. Writes are recorded in the audit log.
Signing in a device
Programs like the LLeMbas CLI sign in with the device-code flow (RFC 8628): the program shows a link and a code, you approve it on the server seeing the device’s name and address, and the program receives a token. /.well-known/lembas.json tells a program it has found a LLeMbas server.
The administrator’s API
| Endpoint | Scope |
|---|---|
GET /api/v1/admin/users, /api/v1/admin/users/{id} | admin:read |
GET /api/v1/admin/usage?period=YYYY-MM | admin:read |
GET /api/v1/admin/audit (filters: action, actor, since, until; format=csv) | admin:audit |
| Credits and plans — see Credits | admin:credits |