LLeMbas Server

A home for your models

A self-hosted web UI and API for local and remote language models. Chats, agent chats, a library, reports, voice, images, people and permissions — on your own machine.

Features

Everything below is in 1.0.0. Each part can be switched on, off or limited by an administrator.

Chat

  • Streaming replies with live Markdown and server-side syntax highlighting
  • Reasoning shown in its own collapsible block, with how long it took
  • Stop and keep what arrived; edit an earlier message and run on from there
  • Replies keep running when you leave the page — a dot and a notification say when they land
  • Attachments: images for vision models, PDFs and text files extracted into the prompt
  • @ to name a document, / for commands, reasoning effort per chat
  • System prompts per instance, per model and per chat; folders; search across everything

Agent chats

  • A chat in the browser whose work runs on a machine with the LLeMbas CLI
  • Manual, Edit, Auto and Plan modes, enforced on every step
  • Approvals, questions and plans as cards — answer here or in the terminal
  • A real terminal panel beside the chat, on the same machine
  • Sessions started in the terminal show up as chats, and the other way round

Library

  • Knowledge: documents, images and web pages in named bases
  • Notes a model writes and finds again later
  • Memory: short facts about you, in front of the model on every turn
  • Skills: saved procedures a model follows — and writes
  • Keyword search, plus meaning when you pick an embedding model

Tools and automation

  • Web search the model calls when it needs it: DuckDuckGo, SearXNG or Firecrawl
  • Your own HTTP tools, and MCP servers by URL
  • Helpers: a reply hands self-contained work to sub-agents that run in parallel
  • Schedules in plain words — “every Monday at nine” — that file reports or start work
  • Web push, so a report at seven in the morning reaches a closed browser

Images, prototypes, voice

  • Image generation through ComfyUI, with your workflow, defaults and an optional review pass
  • Prototypes: a live HTML page in the reply, sealed off and with no network unless you allow it
  • Dictation and read-aloud against any OpenAI-compatible audio endpoint; each person picks a voice

People and control

  • Users and groups; permissions that add up across groups
  • Quotas: tokens a month, replies at once, agent time, images a day, helpers a reply
  • Credits with plans and renewals, set by administrators or over the API
  • Read-only sharing of documents, notes, skills and reports with people and groups
  • Data groups: which provider’s models may read which part of your data
  • An audit log, and a page that answers “what can this account actually do?”

An API

  • OpenAI-compatible /v1: models, chat, embeddings, audio, images, usage
  • Tokens with scopes, checked against their holder on every request
  • Your library over MCP at /mcp
  • API-only mode: models behind one key, with the product switched off

Everywhere

  • Installable as an app on phones and desktops
  • Nine languages: English, Slovak, Czech, German, Polish, Hungarian, French, Spanish, Italian
  • Two themes — Moria (dark) and Shire (light)
  • No build step and nothing fetched at runtime

What you need

  • Docker, or Python 3.11+ on Linux.
  • HTTPS in front for anything beyond localhost. Browsers only allow the microphone and installing the app over HTTPS or on localhost, so a plain-http install on a LAN address cannot dictate or be installed.
  • At least one model endpoint — a local runner such as llama.cpp, LM Studio or Ollama, or an account at a provider.

One worker. LLeMbas runs as one process, on purpose. Replies in flight, terminal sessions and the schedule ticker live in it, so two replicas would fire every schedule twice.

Install

Docker

The repository carries a Dockerfile and a docker-compose.yml. The image runs as a non-root user, keeps its data in a named volume on /data, and bakes in no secret key and no data.

shell
git clone https://github.com/LLeMbas/LLeMbas.git
cd LLeMbas

# A secret key, generated once and kept: it signs sessions and encrypts stored API keys
echo "LEMBAS_SECRET_KEY=$(python3 -c 'import secrets; print(secrets.token_urlsafe(48))')" > .env

docker compose up -d --build      # serves on http://127.0.0.1:8080

The compose file publishes on 127.0.0.1:8080 and expects a TLS reverse proxy (Caddy, nginx, Traefik…) on the same host. docker compose down -v is the command that deletes the data volume; down alone keeps it.

From source, to try it

shell
git clone https://github.com/LLeMbas/LLeMbas.git
cd LLeMbas
python3 -m venv .venv && . .venv/bin/activate
pip install -e ".[search]"        # search: DuckDuckGo web search; drop it if you do not want it

cp .env.example .env
lembas-server secret-key          # paste the result into LEMBAS_SECRET_KEY in .env
lembas-server serve               # http://127.0.0.1:8080

A machine of its own

For a permanent install, deploy/ holds a systemd unit, an nginx virtual host and an idempotent installer. It creates a service account, clones the repository, builds a virtualenv, generates the environment file with a fresh secret key, installs the unit and the vhost with a self-signed certificate, and enables the service.

shell
SITE_HOST=chat.example ./deploy/install.sh

Every setting can be overridden from the environment — the host name, port, service user, install prefix, branch, update channel and whether to install the update helper. The full table is in deploy/README.md.

On Proxmox, deploy/lxc-install.sh creates an unprivileged container and runs the same installer inside it, with the update helper switched on:

shell
CTID=<container id> SITE_HOST=chat.example ./deploy/lxc-install.sh

First steps

  1. Open the address and create the first account. It becomes the administrator.
  2. Go to Admin → Connections and add an endpoint — for a local runner usually http://localhost:1234/v1 with no key. Press Test & refresh.
  3. Enable the models you want under Admin → Models, and mark which ones take tools, reasoning or vision.
  4. Start a chat. To let a model search the web, configure Admin → Web search.
  5. For agent chats, turn on Admin → Agents and link a machine with the LLeMbas CLI.

More in Connections and Agent chats.

Updating

A container

The image is the unit of deployment, so updating is building a new one from the newer source. There is deliberately no update button inside a container: it would need the Docker socket, which is root on the host.

shell
git pull
docker compose up -d --build

A machine of its own

Admin → Updates shows the version running, what is available on the host’s channel, the release notes and the commits between. Two channels:

ChannelFollowsFor
stable (default)the newest release tag vX.Y.Zanybody running it
edgethe tip of the configured branchwhoever is building it

The Update button is opt-in: install with INSTALL_UPDATE_HELPER=1 and a small root-owned helper applies the update when an administrator asks. It always deploys the host’s configured channel and nothing else. Without it, the page prints the command to run by hand.

shell
INSTALL_UPDATE_HELPER=1 SITE_HOST=chat.example ./deploy/install.sh

Take a backup first whenever you like: lembas-server backup makes a checked copy of the database while it runs.

Configuration essentials

Settings are environment variables prefixed LEMBAS_, usually kept in .env (.env.example is annotated). Almost everything else is set in the web interface under Admin.

VariableDefaultPurpose
LEMBAS_SECRET_KEYnoneSigns sessions and encrypts stored API keys. Set it and keep it: changing it signs everyone out and makes stored keys unreadable.
LEMBAS_DATA_DIR./dataThe SQLite database and uploads.
LEMBAS_HOST / LEMBAS_PORT127.0.0.1 / 8080Bind address.
LEMBAS_ALLOW_SIGNUPtrueWhether people may register themselves — the initial value only; Admin → General wins once saved. The first account is always an administrator.
LEMBAS_DEFAULT_THEMEmoriamoria (dark) or shire (light) for signed-out visitors.
LEMBAS_SESSION_TTL2592000Session lifetime in seconds (30 days).
LEMBAS_REQUEST_TIMEOUT300Seconds to wait on an upstream model.
LEMBAS_LOG_LEVELinfodebug, info, warning or error.

Commands

shell
lembas-server serve          # run the server
lembas-server info           # where data lives, what is configured
lembas-server secret-key     # generate a value for LEMBAS_SECRET_KEY
lembas-server create-admin   # create or promote an administrator
lembas-server backup         # a checked copy of the database, while it runs

The server’s command is lembas-server. Plain lembas is the LLeMbas CLI’s command, so both can live on one machine.

Source and licence

MIT licensed. Written in Python — FastAPI, Jinja and htmx, with SQLite — and rendered on the server, so there is no JavaScript build.