Control plane · local AI stack · MIT
Your localhost, on one screen.
There are forty-one open-source AI tools in this catalog, sorted into eight architectural layers. localhost-ai-hub is the dashboard that knows where each of them lives, the gateway that holds the Docker socket and talks to Ollama on your behalf, and the honest answer to “is it actually up?”
No account, no API key, no SaaS dependency in the codebase. Every request it
makes goes to a port on the host you point it at — by default, localhost.
browser :3000 │ /api → vite proxy ▼ gateway :3001 ├─ dockerode → /var/run/docker.sock ├─ http → ollama :11434 ├─ http → vllm :8005 └─ fs → gateway/data/db.json
Four destinations. All of them are on this machine.
- Version
- 2.5.0
- Services
- 41
- Layers
- 8
- Tests green
- 159
- Own ports
- 3000 / 3001
- Cloud services
- 0
Every tool in the catalog, plotted by the port it wants.
Twenty-one of the forty-one default into the 8000s. That is the whole
problem in one column. Two more — SuperAGI and OpenHands — both default to
:3000, which is also where the dashboard itself runs. A test in
catalog.test.ts pins that one known collision, so the next one fails CI instead
of failing on your machine. Per-service port overrides live in Settings; the catalog keeps
each tool's real upstream default rather than inventing a conflict-free one.
Brain Logic Hands Brain Power Knowledge Interface Studio Watchtower
One bar per service, stacked by port within its thousand-band, coloured by layer. Outlined bars share a port with another tool.
The dashboard is the catalog. The gateway is the part that does things.
It holds the Docker socket, reads your hardware, discovers what Ollama has loaded,
and exposes one REST surface that the UI consumes. Here is that surface, endpoint by endpoint
— every path below exists in gateway/routes/.
-
Containers
dockerode, never a shell
- GET /api/services
- POST /api/services/:name/start
- POST /api/services/:name/stop
- POST /api/services/:name/restart
- GET /api/services/:name/logs
Start, stop, restart and tail any container from the dashboard. Names are validated against
/^[a-zA-Z0-9][a-zA-Z0-9_.-]*$/before they reach the daemon; a container name has never touched a shell since v2.0.0. -
Hardware telemetry
polled and streamed
- GET /api/telemetry
- GET /api/telemetry/stream
CPU utilisation, cores and model; system RAM; disk; the Node process’ own memory; GPU and VRAM utilisation and temperature via
nvidia-smi, with a fallback when there is no NVIDIA card. -
Models
Ollama and vLLM, discovered
- GET /api/models
- GET /api/models/active
- POST /api/models/switch
- POST /api/models/preload
Finds what is loaded on
:11434and vLLM, reads parameter sizes, quantisations and families, and keeps a model warm with a configurablekeep_aliveso the next switch has no cold start. -
OpenAI-compatible layer
point your existing client at :3001
- GET /v1/models
- POST /v1/chat/completions
- POST /v1/completions
- POST /v1/embeddings
Chat completions stream over SSE or return whole. Anything that already speaks the OpenAI wire format can be aimed at the gateway instead of at api.openai.com.
-
RAG, offline
no embedding API involved
- POST /api/rag/index
- POST /api/rag/query
- POST /api/rag/search
- GET /api/rag/documents/:id/chunks
Sliding-window chunking on sentence boundaries, 384-dimensional L2-normalised vectors generated locally (with an Ollama embedding fallback), cosine similarity search, and answers synthesised with source citations.
-
Benchmarks
numbers you produce, not numbers we print
- POST /api/benchmarks/run
- GET /api/benchmarks
- GET /api/benchmarks/:id
Measures time-to-first-token, tokens per second, total latency and memory footprint across one model or several, comparatively. The results are yours; this page does not quote any.
-
Agents & workflows
execution, not CRUD theatre
- POST /api/agents/:id/task
- GET /api/agents/:id/tasks
- POST /api/workflows/:id/execute
- GET /api/workflows/:id/runs
Agent tasks run against your local Ollama. Workflows run
prompt,httpanddelaysteps in sequence, a failed step stops the run, and history is capped at 20 tasks and 10 runs per record. -
Event stream
server-sent, with heartbeats
- GET /api/events
A Docker snapshot every five seconds plus task and workflow completion events, so a client can stop polling.
It also tells you when it cannot tell. Health probes run from the browser, so a service with no CORS headers reports CORS blocked rather than pretending to be down, and a TCP-only check reports unknown rather than guessing. That distinction is in the code, in the tests, and on the badge.
Eight layers is not a marketing count. It is the index.
Every tool in the catalog belongs to exactly one layer, and the dashboard's pages, filters and per-layer counts are all derived from that single field. Change the field, the whole app re-sorts itself. Scroll, and the plan view pulls apart.
-
L1The Brain Model runtime7
Local model servers. Everything else eventually calls into this layer.
LM Studio:1234GPT4All:4891vLLM:8005LocalAI:8080llama.cpp Server:8085Ollama:11434SGLang:30000
-
L2The Logic & Glue Automation & orchestration5
Visual workflow engines that wire the other services to each other.
n8n:5678Langflow:7862Activepieces:8000Windmill:8001Dify:8700
-
L3The Hands Browser & web automation3
Agents that drive a real browser: forms, clicks, scraping, replay.
LaVague:7860Skyvern:8882Playwright UI:9323
-
L4The Brain Power Multi-agent frameworks5
Frameworks for running several specialised agents against one task.
SuperAGI:3000OpenHands:3000AutoGen Studio:8081CrewAI:8100Letta:8283
-
L5The Knowledge RAG & vector storage7
Ingestion, embeddings, vector search and document chat.
AnythingLLM:3010Qdrant:6333PrivateGPT:8101Chroma:8200Weaviate:8300Mem0:8400RAGFlow:8600
-
L6The Interface User-facing UIs8
The surfaces you actually type into, plus the API gateways behind them.
AI Gateway (local):3001Flowise:3100Perplexica:3300LiteLLM:4000Text Gen WebUI:7861Tabby:8180Open WebUI:8501SearXNG:8888
-
L7The Studio Image, video & speech4
Generation and transcription on your own GPU.
ComfyUI:8188Kokoro TTS:8880Speaches:8960InvokeAI:9090
-
L8The Watchtower LLM tracing & evals2
Traces, costs, prompt versions and evaluations for the calls above.
Langfuse:3200Arize Phoenix:6006
There is no cloud in it to switch off.
No key, no account
There is no sign-in, no API key field and no third-party SaaS client anywhere in the gateway or the dashboard. The QA pass that went looking for key-gated features found none, which is why every feature was testable offline.
The one outbound rule
A workflow's http step is checked against a hostname set before it runs.
Anything that is not localhost, 127.0.0.1 or ::1
comes back as a failed step with one sentence:
No auth, deliberately
There is no authentication, no rate limiting and no HTTPS, because this binds to your own machine. That is a documented decision, not an oversight — if you expose it beyond localhost, put Caddy or nginx in front of it first.
Forty-one tools, and the port each one answers on.
Ports, docs links, start hints and health-check config for every entry live in one TypeScript file. Filter by layer:
L1 The Brain Model runtime 7
- :1234LM StudioDesktop app to find, download and serve local models, with a built-in server.
- :4891GPT4AllLocal models on any CPU. Chat with your documents fully offline.
- :8005vLLMPagedAttention throughput serving behind an OpenAI-compatible API.
- :8080LocalAIDrop-in OpenAI API replacement running on your own CPU or GPU.
- :8085llama.cpp ServerTiny GGUF server that runs anywhere, including plain CPUs.
- :11434OllamaThe de-facto local LLM server. Llama 3, Mistral, Gemma, 100+ models, one command.
- :30000SGLangRadixAttention serving engine; strong at structured and JSON generation.
L2 The Logic & Glue Automation & orchestration 5
- :5678n8n400+ integrations, AI nodes, visual builder, self-hostable forever.
- :7862LangflowDrag-and-drop LangChain-style flows, servable as an API.
- :8000ActivepiecesOpen-source Zapier alternative with a clean builder.
- :8001WindmillTurns Python, TypeScript, Bash and Go scripts into workflows and APIs.
- :8700DifyLLM app platform: visual chains, agents, RAG pipelines, model management.
L3 The Hands Browser & web automation 3
- :7860LaVagueLarge Action Model agents: describe the goal, it drives the browser.
- :8882SkyvernVision plus LLM browser automation. No brittle selectors.
- :9323Playwright UITest runner with a visual trace viewer for recorded browser sessions.
L4 The Brain Power Multi-agent frameworks 5
- :3000SuperAGIAutonomous agents with tools, memory and a telemetry dashboard.
- :3000OpenHandsAutonomous dev agent: writes code, runs tests, browses docs.
- :8081AutoGen StudioMicrosoft's multi-agent framework with a visual studio UI.
- :8100CrewAIRole-playing agent crews: researcher, writer, critic.
- :8283LettaStateful agents with long-term memory (formerly MemGPT), over REST.
L5 The Knowledge RAG & vector storage 7
- :3010AnythingLLMPrivate document assistant. Ingest PDFs and URLs into workspaces.
- :6333QdrantRust vector search with advanced filtering and payload storage.
- :8101PrivateGPTDocument chat, fully offline, from ingestion through to the answer.
- :8200ChromaAI-native embedding database with a dead-simple Python/JS SDK.
- :8300WeaviateVector database with built-in models and a GraphQL API.
- :8400Mem0Persistent memory layer so agents remember across sessions.
- :8600RAGFlowDeep document understanding: real OCR, table parsing, grounded citations.
L6 The Interface User-facing UIs 8
- :3001AI Gateway (local)This project's own Express gateway — the REST API the dashboard calls.
- :3100FlowiseDrag-and-drop builder for LLM apps and agents.
- :3300PerplexicaSelf-hosted answer engine: SearXNG plus your local model, with citations.
- :4000LiteLLMOne OpenAI-compatible endpoint for 100+ models, with fallbacks.
- :7861Text Gen WebUIThe oobabooga gradio UI: extensions, fine-tuning, notebook mode.
- :8180TabbySelf-hosted coding assistant served from your own GPU.
- :8501Open WebUIThe ChatGPT-style front end for local models. RAG, web search, plugins.
- :8888SearXNGPrivate metasearch. Live web results for agents and RAG, no tracking.
L7 The Studio Image, video & speech 4
- :8188ComfyUINode-graph image and video generation: SD, SDXL and Flux workflows.
- :8880Kokoro TTS50+ voices across 9 languages, streaming, OpenAI-compatible.
- :8960Speaches"Ollama for speech": faster-whisper STT plus Kokoro/Piper TTS, one API.
- :9090InvokeAIImage generation on a clean unified canvas.
L8 The Watchtower LLM tracing & evals 2
- :3200LangfuseSelf-hosted LLM observability: traces, costs, prompt versions, evals.
- :6006Arize PhoenixOpenTelemetry-native tracing for LLM, RAG and agent pipelines.
Clone it, install it, open :3000.
# 1 — get it $ git clone https://github.com/Arnav1771/localhost-ai-hub.git $ cd localhost-ai-hub # 2 — node check, .env files, deps, data dirs $ ./scripts/init.sh # 3 — two terminals $ npm run dev:gateway # :3001 $ npm run dev:dashboard # :3000
$ docker compose up --build -d $ curl localhost:3001/health
- Node18 or newer, npm 9 or newer.
- DockerOptional. Only the container-management endpoints need it — the gateway starts without a daemon and reports 503 on those routes.
- Lockfile
package-lock.jsonis git-ignored on purpose. Usenpm install, notnpm ci. - OllamaOptional, but agent tasks and warm model switching have nothing to
talk to without it. Default
http://localhost:11434.
The unflattering half of this page.
There is a v3 specification for this project that describes a graph runtime, artifacts, evaluations and human-in-the-loop approvals. Almost none of it exists yet. Here is the split, so nothing on this page has to be taken on faith.
In the build
v2.5.0- Catalog41 tools, 8 layers, ports, docs links and start hints, with an integrity test suite
- HealthPer-service browser probes: up, down, CORS-blocked, or honestly unknown
- ContainersStart / stop / restart / logs over the Docker socket
- TelemetryCPU, RAM, disk, process memory, GPU and VRAM — polled or streamed
- ModelsDiscovery across Ollama and vLLM, warm preloading, zero-cold-start switching
- OpenAI API/v1 chat, completions and embeddings, streaming included
- RAGOffline chunking, 384-dim embeddings, cosine search, cited answers
- BenchmarksTTFT, tokens/sec, latency and memory, single or comparative
- ExecutionAgent tasks on Ollama; sequential workflow runs with per-step results
- EventsSSE stream of Docker snapshots and completions
- ShellCommand palette, light and dark themes, per-service port overrides
- Tests159 across 26 suites — 82 Jest on the gateway, 77 Vitest on the dashboard
Specified, not built
v3 spec- Graph runtimeA stateful LangGraph execution engine with conditional routing and retries
- Run inspectorStep-level traces, timelines and replay for a run that has finished
- EvaluationsScored, repeatable evals rather than one-off benchmark rows
- ApprovalsHuman-in-the-loop gates that can pause a run and wait
- ArtifactsFirst-class outputs — files, tables, images — attached to the run that made them
- Scheduler / queueCron triggers and a real job queue instead of request-scoped execution
- SecretsA managed store; today configuration is env vars and a JSON file
- Tool sandboxAn isolation policy for tools that execute code
- SQLiteThe database is a synchronous JSON file today. It should not stay one.
- Streaming tasksAgent output arrives whole; Ollama can stream it token by token
A desktop build — one executable instead of two npm processes — is the
direction, not a release. It has not been started. Today it is npm run dev:gateway
and npm run dev:dashboard, or a two-service compose file, and this page will say so
until that changes.