API

A program can use LLeMbas as the person it acts for: their models through an OpenAI-compatible API, their library over MCP, and — for an administrator — accounts, usage and the audit log.

Tokens

Make a token under Settings → Security → API tokens: a name, an expiry (never, 30 or 90 days, a year) and its scopes. You need the Use the API permission. The value is shown once; afterwards only its first characters, to tell tokens apart. Send it as Authorization: Bearer lmb_….

ScopeAllowsNeeds
models/v1: your models, chat, embeddings, audio, images, usage; search and fetchUse the API
library/mcp: your memory, notes, skills and knowledgeUse the API
linkbeing driven from the server as a linked deviceUse the API
admin:readaccounts and usage, read onlyadministrator
admin:auditthe audit logadministrator, or Read the audit log
admin:creditstop-ups, adjustments, plansadministrator, or Grant credits
  • Scopes are checked against the holder on every request. Take Use the API away from someone and all their tokens stop at once.
  • A token is never accepted on a web page, and a browser session is never accepted on the API.
  • Each token may make a limited number of requests a minute (120 by default, set in Admin → General); over it the answer is 429 with Retry-After.
  • Your devices and tokens are listed under Settings; an administrator sees everyone’s on each user’s page.

Endpoints

EndpointWhat
GET /v1/modelsThe models you could start a chat with, as provider/model, each with a kind: chat, embedding, stt, tts or image — plus context, output limit, efforts, vision and tools.
POST /v1/chat/completionsOpenAI’s chat completions, streamed or not.
POST /v1/embeddingsEmbeddings from an embedding model.
POST /v1/audio/transcriptionsSpeech to text: a multipart file, optionally language. Uses the server’s transcription model.
POST /v1/audio/speechText to speech: input, optionally voice, speed, response_format (mp3, wav, flac, aac).
POST /v1/images/generationsImage generation through the server’s ComfyUI.
GET /v1/usageThis month’s tokens and credits, your plan and balance, and this token’s share.
POST /api/v1/search{"query": …, "max_results": n} → results with title, url and snippet.
POST /api/v1/fetch{"url": …} → the page’s title and text.
GET /api/v1/instanceThe server’s name, your account, your default model, and which services you may use.
POST /mcpYour library over MCP (streamable HTTP), with a library token.

The API is stateless: nothing is kept, no chat and no message. Each request goes to the model’s connection with that connection’s own key, which never leaves the server. Requests are counted and limited like replies in the interface — tokens a month, replies at once, credits — and errors come back in OpenAI’s shape. An endpoint whose feature is switched off for you answers 404.

Examples

shell
# list the models you may use
curl -s https://your-server/v1/models -H "Authorization: Bearer $TOKEN"

# chat, streamed
curl -s https://your-server/v1/chat/completions \
  -H "Authorization: Bearer $TOKEN" -H "Content-Type: application/json" \
  -d '{"model": "llama/qwen", "messages": [{"role": "user", "content": "Hello"}], "stream": true}'

# embeddings
curl -s https://your-server/v1/embeddings \
  -H "Authorization: Bearer $TOKEN" -H "Content-Type: application/json" \
  -d '{"model": "llama/bge-m3", "input": "waybread"}'

# transcribe, then speak
curl -s https://your-server/v1/audio/transcriptions \
  -H "Authorization: Bearer $TOKEN" -F file=@note.wav
curl -s https://your-server/v1/audio/speech \
  -H "Authorization: Bearer $TOKEN" -H "Content-Type: application/json" \
  -d '{"input": "Speak, friend, and enter."}' -o reply.mp3

# draw
curl -s https://your-server/v1/images/generations \
  -H "Authorization: Bearer $TOKEN" -H "Content-Type: application/json" \
  -d '{"prompt": "a mallorn leaf, pixel art"}'

# this month's usage
curl -s https://your-server/v1/usage -H "Authorization: Bearer $TOKEN"

Any OpenAI client library works: set its base URL to https://your-server/v1 and its API key to your token.

python
from openai import OpenAI

client = OpenAI(base_url="https://your-server/v1", api_key="lmb_…")
reply = client.chat.completions.create(
    model="llama/qwen",
    messages=[{"role": "user", "content": "Hello"}],
)
print(reply.choices[0].message.content)

Your library over MCP

/mcp serves the same library tools a chat has — memory, notes, skills, knowledge — as far as your permissions allow, to any MCP client with a library token. One JSON-RPC message per POST, answered with JSON. Writes are recorded in the audit log.

Signing in a device

Programs like the LLeMbas CLI sign in with the device-code flow (RFC 8628): the program shows a link and a code, you approve it on the server seeing the device’s name and address, and the program receives a token. /.well-known/lembas.json tells a program it has found a LLeMbas server.

The administrator’s API

EndpointScope
GET /api/v1/admin/users, /api/v1/admin/users/{id}admin:read
GET /api/v1/admin/usage?period=YYYY-MMadmin:read
GET /api/v1/admin/audit (filters: action, actor, since, until; format=csv)admin:audit
Credits and plans — see Creditsadmin:credits