Nexus.
The control plane. Phron runs on every machine.
Two parts, one system. Nexus is where your team logs in, your admins manage the fleet, and every request gets routed. Phron is the quiet agent that runs on each machine with a GPU, installs what it's told to, and serves the model. Neither one talks to the outside world unless you decide it should.
One system, two parts, no exceptions.
Nexus is where people work. Phron is where the model runs. Every request crosses that boundary the same way, every time.
Phronexus is not a third layer sitting on top — it is Nexus and Phron running together. The control plane and the node agent, fused into one product, still entirely on your side of the wall.
A · The two-part architecture
Nexus orchestrates. Phron executes.
Your team, your admins, and any connected app all talk to Nexus — the web console and API. Phron sits quietly on each machine with a GPU, waiting for instructions. It never receives a request directly from a person. It only ever hears from Nexus.
B · The request, step by step
Six steps, every time.
A user picks a model. Nexus checks who they are, what they're allowed to see, and which machine has that model ready. It sends the job over an already-open connection. The machine runs it and streams the answer back the same way it came.
01
Pick a model
Only loaded, allowed models appear
02
Nexus checks
Identity, permissions, quota
03
Nexus routes
Finds the node hosting that model
04
Phron receives
Over an already-open connection
05
The model runs
llama.cpp on that machine
06
Answer streams back
Through Nexus, to the user
Your GPU never has to face the internet.
Every machine in your fleet connects outward, the same way a build agent registers with a CI server. Nothing external ever needs a route in.
A · Outbound only
The connection only ever goes one way.
Phron opens a connection to Nexus and keeps it open. It never listens for the outside world. You can run your entire fleet behind a firewall with no inbound rules for inference at all — the same posture as a laptop that only ever calls out, never gets called.
No inbound port. No exposed GPU.
B · Who can prove what
Every credential has one job.
People sign into Nexus with their own account. Integrations use a scoped API key, never a person's password. Each machine gets its own enrollment credential, generated once, revocable any time. Nobody downstream ever sees a machine's credentials or its direct address.
| Who | How they connect | What they hold |
|---|---|---|
| Users | Sign in to Nexus | Password + session |
| Integrations | Call the API | Scoped API key (nxs_live_…) |
| Machines | Register once with Phron | One-time enrollment token |
C · What Nexus enforces on every request
Nothing reaches a model by accident.
Authentication, role, model allowlist, and quota — checked before a single token is generated. A user only ever sees models an admin has explicitly published to them, and only if that model is actually loaded and ready.
01
Authenticated
Signed-in user or valid API key
02
Authorised (role)
Admin, User, or Restricted
03
Allowlisted (model)
Explicitly published to this seat
04
Within quota
Tokens and rate still available
Every machine you own, as one fleet.
Add a node once. From then on it's a line in a list — its status, its models, its uptime — until you decide otherwise.
A · The control plane, live
From empty fleet to a monitored node — the real screens.
Watch the path an admin actually takes: register a machine, install Phron, enroll with a one-time token, then open the node for overview, updates, and live hardware telemetry. Nothing here is a stock illustration. Hover to pause; click a step to jump.
LLM NODES
Register Phron machines, pause to free GPU, resume to reload the last models.
- hq-inference-01online
Connected · hq-inference-01
Loaded 2 / 2Last seen Just nowmistral-smallqwen3-14b - office-milanonlinepaused
Connected · office-milan
Loaded 0 / 1Last seen 12m ago - Local Serverpending
Awaiting enroll · local-server
Loaded 0 / 0Last seen — - lab-aoffline
Node not connected · lab-a
Node not connected
Loaded 0 / 1Last seen 9h ago
B · Enrollment
One token, one command.
Create a pending node in Nexus and you get a one-time token. Run one command on the target machine and it registers, opens its connection, and appears online — usually in under a minute.
01
Name the node
Pick the Nexus API URL Phron will call
02
Get the token
One-time enrollment credential
03
Install & enroll
phron enroll --url … --token …
04
It comes online
Outbound WebSocket — no inbound port
C · The fleet list
Status, at a glance.
Every node's state is visible from the moment it registers: pending, online, offline, or paused. Click into any of them for overview, updates, models, and live metrics.
| Node | State | What you see |
|---|---|---|
| hq-inference-01 | Online | Models loaded · last seen just now |
| office-milan | Paused | Config kept · GPU freed |
| Local Server | Pending | Waiting for enroll |
| lab-a | Offline | Last seen hours ago |
D · Node lifecycle
Pause without losing configuration.
Idle a node for maintenance and its models unload cleanly — but Nexus remembers what it was running, so resuming reloads the same setup automatically. Decommissioning removes it and its allowlist entries in one action.
01
Pending
Token issued, waiting for enroll
02
Online
Outbound WebSocket open
03
Paused
Models unloaded, config kept
04
Online again
Same setup reloads automatically
Install what fits. Load what you need.
An engine is the runtime; a model is the weights. Nexus keeps the catalog, verifies every download, and tells you honestly what your hardware can actually run.
A · Engines & models, live
From package install to a loaded model — the real screens.
One-click engines on Phron, a filtered catalog with VRAM fit, a recommendation wizard, then load / unload with progress and advanced settings that warn before you OOM. Hover to pause; click a step to jump.
Local Server
onlineConnected · gpu-rack-03 · 0.1.2.20240825 · local-server
Install engine packages
Required before loading models (e.g. llama.cpp-cpu / cuda)
GGUF runtime — NVIDIA CUDA 12.x · chat & coding weights
CPU-only fallback — AVX2 baseline, no GPU required
High-throughput OpenAI-compatible server — continuous batching
Image generation node — Stable Diffusion / FLUX workflows
Speech-to-text — multilingual transcription on GPU
B · Engines
Runtime first. One click on the node.
An engine is the runtime that can actually run weights — GGUF chat, high-throughput serving, image generation, speech. Install the package that matches the machine before you touch the catalog. Nexus ships curated builds; Phron places them on disk.
| Package | When you pick it | Install |
|---|---|---|
| llama.cpp / cuda12 | GGUF chat & coding on NVIDIA | One click · becomes active |
| llama.cpp / cpu | No GPU · AVX2 fallback | One click on Setup |
| vllm / cuda12 | High-throughput OpenAI-compatible | One click · progress live |
| comfyui / cuda12 | Image generation · SD / FLUX | One click on Setup |
| whisper / cuda12 | Speech-to-text on GPU | One click · becomes active |
C · Catalog & fit
Filter by what the card can hold.
The catalog is curated GGUFs with verified checksums. Nexus scores fit against the node’s VRAM, marks what is comfortable or too heavy, and can recommend from concurrent users, speed, context, and capabilities — then install with a live progress bar.
| Signal | What it means |
|---|---|
| VRAM tier filters | CPU · 8 GB · 12 GB · 24 GB · … |
| Recommended strip | Best fits for this GPU, one-click |
| TOO HEAVY | Would overfill VRAM — install still allowed, clearly marked |
| Help me choose | Wizard → ranked shortlist for this node |
| Download progress | Bytes on the wire · cancel anytime |
D · Load & advanced
Easy load. Honest warnings.
Load and unload from the node’s Models tab. Advanced load remembers engine, context, GPU layers, and presets — and surfaces OOM risk with one-click safer settings before you commit.
01
Pick a model on disk
Detail · Load · or Advanced
02
Watch progress
Engine + layers coming up · Cancel if needed
03
Unload when idle
Frees VRAM · config kept for next load
04
Advanced if you must
Context, slots, fit-to-VRAM — with OOM recommendations
The interface your team already knows.
Streaming answers, editable messages, saved conversations — the shape of every AI tool your staff already use, running entirely on your own models.
A · Portal chat, live
Messages, tools, voice, files — the real portal.
Watch a conversation unfold: vision on an upload, a voice note, a generated PDF, then an agent turn that writes code and an Imagine cover. API keys stay in the portal for chat and clients. Hover to pause; click a step to jump.
Summarize the attached site photo for the ops standup — risks only.
B · Chat & agent tools
The interface your team already knows — with tools that act.
Streaming answers, editable threads, attachments, and voice. Agents can turn on Web, Vision, and Imagine, emit code and files, and keep token usage visible on every reply.
| Capability | What users see |
|---|---|
| Streaming chat | Live replies · token counts · model picker |
| Vision & voice | Image uploads · mic notes into the thread |
| Imagine | Describe → generated image on the node |
| Code & files | Scripts and PDFs attached to the answer |
| Web / Params | Optional tools and generation knobs per turn |
C · API keys
Create once. Trace every call.
Keys authenticate Nexus Chat and any client that hits the OpenAI-compatible API. The secret is shown once at creation; last-used timestamps stay on the list so you know what is still live. Wire those keys into external tools in Integrations.
01
Label the key
e.g. cursor-integration
02
Create & copy
nxs_live_… shown once
03
Select in chat
Or paste into an external client
04
Follow usage
Last used · revoke anytime
One organisation. Clear roles. Groups that stick.
An owner sits at the top. Admins manage seats and limits. Users sign in, join permission groups, and work under the quotas you set — invite links or direct Add User.
A · Control directory, live
Invite, group, and seat every person.
The real User Management screen: invite links, groups, the directory — then New user, and the User Detail drawer (Overview, API keys, Usage). Hover to pause; click a step to jump.
USER MANAGEMENT
Invite links set role and limits; users choose their profile at signup. Direct Add User creates an active account immediately.
Invite links
Create a reusable signup code for northwind ops. Set role and limits here — the user picks their name and email when they sign up.
No active invites. Create one to share a signup link.
Need to set the password yourself? Use Add User in the table below.
- Alex Ortega·just now
Permission groups
Select users in the table, then add them to a group. Models are allowed for everyone unless you restrict a model on a node.
- ops-core2 members · LLM on
- vision-pilot1 member · LLM on
| USER | ROLE | STATUS | TOKEN LIMIT | PROMPT LIMIT | LAST ACTIVE | ACTIONS | |
|---|---|---|---|---|---|---|---|
AO Alex Ortega alex@northwind.ops | Owner | active | 500K | 10,000 | just now | ||
MC Maya Chen maya@northwind.ops | Admin | active | 250K | 8,192 | 2h ago | ||
JL Jordan Lee jordan@northwind.ops | User | active | 100K | 4,096 | yesterday |
B · Organisation structure
Owner → admins → users → groups.
northwind ops is one organisation. The owner owns the tenant. Admins manage seats, invites, and limits. Users sign in under those limits. Permission groups (ops-core, vision-pilot) batch LLM access so you do not edit every seat by hand.
| Layer | What it does |
|---|---|
| Organisation | Tenant boundary · name · invite codes |
| Owner | Full control · cannot be demoted |
| Admin | Users, groups, limits, invites |
| User / Restricted | Portal access under quota |
| Permission groups | Batch LLM access · assign from the table |
| User detail | Overview · API keys · usage log per seat |
C · How seats get created
Invite link — or Add User with a password.
Invite links bake role and limits into a signup code; the person chooses name and email. Add User creates an active account immediately when you need to set the password yourself.
01
Create invite or Add User
Role · token · prompt · rate limits
02
They sign in
Portal chat under those quotas
03
Optional: group them
Select rows → Add to ops-core
04
Revoke or suspend anytime
Directory actions · audit logged
Your seat — or the whole organisation.
Portal Usage is personal: quota, recent requests, and drill-down by API key. Control Analytics is org-wide: tokens, cost, top users, and traffic across every seat.
A · Two scopes
Your seat — or the whole organisation.
Portal Usage is personal. Control Analytics is organisation-wide. They never mix in the same screen — each person sees their own quota and keys; admins see the fleet.
| Surface | Audience | Shows |
|---|---|---|
| Portal · Usage | Each user | Quota · requests · usage by key |
| Control · Analytics | Owner / admin | Tokens · cost · top users · traffic |
B · Per user (Portal)
Your usage & groups · requests · by key.
Each person sees their own monthly quota, permission groups, a filterable request history, and charts scoped to a single API key — without anyone else's traffic.
Usage
Quota, charts, and request history
Your usage & groups
Monthly quota across all API keys, permission groups, and activity over the last 30 days.
No permission groups assigned
Groups control model access and LLM permissions. Contact an admin to change membership.
C · Organisation (Control)
Fleet-wide usage analytics.
Admins open Usage Analytics for the rolling window: total tokens, estimated cost, request volume, model mix, and a top-users table.
USAGE ANALYTICS
Live token consumption · model breakdown · top users · rolling window · UTC
| USER | TOKENS | COST | REQUESTS | TOP MODEL | |
|---|---|---|---|---|---|
| Alex Ortega | alex@northwind.ops | 4.8K | $0.01 | 18 | Qwen2.5-VL-7B-Instruct-Q4_K_M.gguf |
| Maya Chen | maya@northwind.ops | 3.1K | $0.01 | 12 | Llama-3.1-8B-Instruct-Q5_K_M.gguf |
| Jordan Lee | jordan@northwind.ops | 2.4K | <$0.01 | 9 | Qwen2.5-Coder-14B-Instruct-Q4_K_M.gguf |
Cost figures are estimated from token volume until per-model billing rates are configured.
Your existing tools, pointed at your own models.
An OpenAI-compatible API means anything already built for OpenAI's format works here with a one-line change: the URL.
A · OpenAI-compatible API
One Base URL. Your existing tools.
Paste /v1, a Nexus key, and a published model id. Same allowlists and quotas as the portal — Cursor, Open WebUI, Continue, and anything else that speaks OpenAI’s format work without a rewrite.
OpenAI-compatible API
Connect Open WebUI, Continue, Cursor, or any OpenAI client to your allowlisted local models through Nexus. Every call uses an API key so usage stays fully traceable.
WSL and other machines cannot use 127.0.0.1 — pick the LAN / vEthernet address.
Authorization: Bearer · or X-Api-Key
- Settings → Connections → OpenAI → add connection
- API URL: paste the Base URL above (must end with /v1)
- API key: your Nexus key (starts with nxs_live_)
- Pick a model from the list below in the client
curl http://10.20.14.8:3001/v1/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer YOUR_NXS_KEY" \
-d '{"model":"LocalServer_Qwen2.5-VL-7B-Instruct","messages":[{"role":"user","content":"Hello"}]}'B · Open source & private clients
From LibreChat to Cursor — same endpoint.
Open-source frontends and private IDEs both connect as OpenAI clients. Nexus does not care which logo is on the window; it cares that the key is valid and the model is allowlisted.
Clients
Same /v1 endpoint — open source and private tools
Point any OpenAI-format client at Nexus. Allowlists and quotas still apply.
Open source
Private & commercial
C · Connect in four steps
Base URL, key, model — done.
Create a key in the portal, copy the Base URL ending in /v1, pick a published model id, and drop them into the client’s OpenAI connection settings.
01
Create a key
Portal → API keys · nxs_live_…
02
Copy Base URL
…/v1 — reachable from the client
03
Paste into the app
OpenAI connection · Bearer auth
04
Select the model id
From the published list on Nexus
Five ways to run Nexus.
Whatever the shape of your business.
See it runningon your own hardware.
Twenty minutes. We look at how your team works, tell you what it would cost, and tell you honestly if you should wait.
NO DECK · NO PITCH · NO FOLLOW-UP SEQUENCE