// THE PHRONEXUS PRODUCT

Nexus.

The control plane. Phron runs on every machine.

Two parts, one system. Nexus is where your team logs in, your admins manage the fleet, and every request gets routed. Phron is the quiet agent that runs on each machine with a GPU, installs what it's told to, and serves the model. Neither one talks to the outside world unless you decide it should.

Phronhq-01PhronmilanPhronlab-aPhronedge-2PhronofficePhronrack-3NEXUSCONTROL PLANE · NODE AGENTS
// SECTION 01 — HOW IT WORKS

One system, two parts, no exceptions.

Nexus is where people work. Phron is where the model runs. Every request crosses that boundary the same way, every time.

Phronexus is not a third layer sitting on top — it is Nexus and Phron running together. The control plane and the node agent, fused into one product, still entirely on your side of the wall.

YOUR ENVIRONMENT · BOTH LOCALPhronnode agentLOCALNexuscontrol planeLOCALPHRONEXUSTHE PRODUCTTWO PARTS · ONE PRODUCT · STILL LOCAL

A · The two-part architecture

Nexus orchestrates. Phron executes.

Your team, your admins, and any connected app all talk to Nexus — the web console and API. Phron sits quietly on each machine with a GPU, waiting for instructions. It never receives a request directly from a person. It only ever hears from Nexus.

USERS & INTEGRATIONSBrowserOpenAI-compatclientsNEXUS CONTROLWebAPIToolsAgentsDatabaseYOUR FLEETPhronllama.cppPhronllama.cppPhronllama.cppRequest flows left → right. Phron never faces the user.

B · The request, step by step

Six steps, every time.

A user picks a model. Nexus checks who they are, what they're allowed to see, and which machine has that model ready. It sends the job over an already-open connection. The machine runs it and streams the answer back the same way it came.

  1. 01

    Pick a model

    Only loaded, allowed models appear

  2. 02

    Nexus checks

    Identity, permissions, quota

  3. 03

    Nexus routes

    Finds the node hosting that model

  4. 04

    Phron receives

    Over an already-open connection

  5. 05

    The model runs

    llama.cpp on that machine

  6. 06

    Answer streams back

    Through Nexus, to the user

// SECTION 02 — SECURITY BY DESIGN

Your GPU never has to face the internet.

Every machine in your fleet connects outward, the same way a build agent registers with a CI server. Nothing external ever needs a route in.

A · Outbound only

The connection only ever goes one way.

Phron opens a connection to Nexus and keeps it open. It never listens for the outside world. You can run your entire fleet behind a firewall with no inbound rules for inference at all — the same posture as a laptop that only ever calls out, never gets called.

YOUR NETWORKPhronnodePhronnodePhronnodeNEXUSon-prem or hostedOUTBOUNDINBOUND BLOCKED

No inbound port. No exposed GPU.

B · Who can prove what

Every credential has one job.

People sign into Nexus with their own account. Integrations use a scoped API key, never a person's password. Each machine gets its own enrollment credential, generated once, revocable any time. Nobody downstream ever sees a machine's credentials or its direct address.

WhoHow they connectWhat they hold
UsersSign in to NexusPassword + session
IntegrationsCall the APIScoped API key (nxs_live_…)
MachinesRegister once with PhronOne-time enrollment token

C · What Nexus enforces on every request

Nothing reaches a model by accident.

Authentication, role, model allowlist, and quota — checked before a single token is generated. A user only ever sees models an admin has explicitly published to them, and only if that model is actually loaded and ready.

  1. 01

    Authenticated

    Signed-in user or valid API key

  2. 02

    Authorised (role)

    Admin, User, or Restricted

  3. 03

    Allowlisted (model)

    Explicitly published to this seat

  4. 04

    Within quota

    Tokens and rate still available

// SECTION 03 — FLEET & NODES

Every machine you own, as one fleet.

Add a node once. From then on it's a line in a list — its status, its models, its uptime — until you decide otherwise.

A · The control plane, live

From empty fleet to a monitored node — the real screens.

Watch the path an admin actually takes: register a machine, install Phron, enroll with a one-time token, then open the node for overview, updates, and live hardware telemetry. Nothing here is a stock illustration. Hover to pause; click a step to jump.

phronexus control · fleet
LLM LIVE · 10:14:54
AO
// FLEET · NODES

LLM NODES

Register Phron machines, pause to free GPU, resume to reload the last models.

4 total1 online1 awaiting registration
  • hq-inference-01online

    Connected · hq-inference-01

    Loaded 2 / 2Last seen Just now
    mistral-smallqwen3-14b
  • office-milanonlinepaused

    Connected · office-milan

    Loaded 0 / 1Last seen 12m ago
  • Local Serverpending

    Awaiting enroll · local-server

    Loaded 0 / 0Last seen —
  • lab-aoffline

    Node not connected · lab-a

    Node not connected

    Loaded 0 / 1Last seen 9h ago

B · Enrollment

One token, one command.

Create a pending node in Nexus and you get a one-time token. Run one command on the target machine and it registers, opens its connection, and appears online — usually in under a minute.

  1. 01

    Name the node

    Pick the Nexus API URL Phron will call

  2. 02

    Get the token

    One-time enrollment credential

  3. 03

    Install & enroll

    phron enroll --url … --token …

  4. 04

    It comes online

    Outbound WebSocket — no inbound port

C · The fleet list

Status, at a glance.

Every node's state is visible from the moment it registers: pending, online, offline, or paused. Click into any of them for overview, updates, models, and live metrics.

NodeStateWhat you see
hq-inference-01OnlineModels loaded · last seen just now
office-milanPausedConfig kept · GPU freed
Local ServerPendingWaiting for enroll
lab-aOfflineLast seen hours ago

D · Node lifecycle

Pause without losing configuration.

Idle a node for maintenance and its models unload cleanly — but Nexus remembers what it was running, so resuming reloads the same setup automatically. Decommissioning removes it and its allowlist entries in one action.

  1. 01

    Pending

    Token issued, waiting for enroll

  2. 02

    Online

    Outbound WebSocket open

  3. 03

    Paused

    Models unloaded, config kept

  4. 04

    Online again

    Same setup reloads automatically

// SECTION 04 — ENGINES & MODELS

Install what fits. Load what you need.

An engine is the runtime; a model is the weights. Nexus keeps the catalog, verifies every download, and tells you honestly what your hardware can actually run.

A · Engines & models, live

From package install to a loaded model — the real screens.

One-click engines on Phron, a filtered catalog with VRAM fit, a recommendation wizard, then load / unload with progress and advanced settings that warn before you OOM. Hover to pause; click a step to jump.

phronexus control · engines
LLM LIVE · 11:24:09
AO
// NODE · FLEET

Local Server

online

Connected · gpu-rack-03 · 0.1.2.20240825 · local-server

// ENGINES

Install engine packages

Required before loading models (e.g. llama.cpp-cpu / cuda)

Models & catalog →
Linux-x86_64 · /opt/phron
llama.cpp / cuda12

GGUF runtime — NVIDIA CUDA 12.x · chat & coding weights

INSTALLED
b78385latestactive
llama.cpp / cpu

CPU-only fallback — AVX2 baseline, no GPU required

b78385latestcatalog
vllm / cuda12

High-throughput OpenAI-compatible server — continuous batching

v0.8.5latestcatalog
download · engine package54%
comfyui / cuda12

Image generation node — Stable Diffusion / FLUX workflows

v0.3.26latestcatalog
whisper / cuda12

Speech-to-text — multilingual transcription on GPU

INSTALLED
v1.7.1latestactive

B · Engines

Runtime first. One click on the node.

An engine is the runtime that can actually run weights — GGUF chat, high-throughput serving, image generation, speech. Install the package that matches the machine before you touch the catalog. Nexus ships curated builds; Phron places them on disk.

PackageWhen you pick itInstall
llama.cpp / cuda12GGUF chat & coding on NVIDIAOne click · becomes active
llama.cpp / cpuNo GPU · AVX2 fallbackOne click on Setup
vllm / cuda12High-throughput OpenAI-compatibleOne click · progress live
comfyui / cuda12Image generation · SD / FLUXOne click on Setup
whisper / cuda12Speech-to-text on GPUOne click · becomes active

C · Catalog & fit

Filter by what the card can hold.

The catalog is curated GGUFs with verified checksums. Nexus scores fit against the node’s VRAM, marks what is comfortable or too heavy, and can recommend from concurrent users, speed, context, and capabilities — then install with a live progress bar.

SignalWhat it means
VRAM tier filtersCPU · 8 GB · 12 GB · 24 GB · …
Recommended stripBest fits for this GPU, one-click
TOO HEAVYWould overfill VRAM — install still allowed, clearly marked
Help me chooseWizard → ranked shortlist for this node
Download progressBytes on the wire · cancel anytime

D · Load & advanced

Easy load. Honest warnings.

Load and unload from the node’s Models tab. Advanced load remembers engine, context, GPU layers, and presets — and surfaces OOM risk with one-click safer settings before you commit.

  1. 01

    Pick a model on disk

    Detail · Load · or Advanced

  2. 02

    Watch progress

    Engine + layers coming up · Cancel if needed

  3. 03

    Unload when idle

    Frees VRAM · config kept for next load

  4. 04

    Advanced if you must

    Context, slots, fit-to-VRAM — with OOM recommendations

// SECTION 05 — CHAT & TOOLS

The interface your team already knows.

Streaming answers, editable messages, saved conversations — the shape of every AI tool your staff already use, running entirely on your own models.

A · Portal chat, live

Messages, tools, voice, files — the real portal.

Watch a conversation unfold: vision on an upload, a voice note, a generated PDF, then an agent turn that writes code and an Imagine cover. API keys stay in the portal for chat and clients. Hover to pause; click a step to jump.

phronexus portal · chat
Q3 board brief outline
Qwen2.5-VL-7B
dock-aisle-b.jpg

Summarize the attached site photo for the ops standup — risks only.

WebVisionImagineParams
Message…

B · Chat & agent tools

The interface your team already knows — with tools that act.

Streaming answers, editable threads, attachments, and voice. Agents can turn on Web, Vision, and Imagine, emit code and files, and keep token usage visible on every reply.

CapabilityWhat users see
Streaming chatLive replies · token counts · model picker
Vision & voiceImage uploads · mic notes into the thread
ImagineDescribe → generated image on the node
Code & filesScripts and PDFs attached to the answer
Web / ParamsOptional tools and generation knobs per turn

C · API keys

Create once. Trace every call.

Keys authenticate Nexus Chat and any client that hits the OpenAI-compatible API. The secret is shown once at creation; last-used timestamps stay on the list so you know what is still live. Wire those keys into external tools in Integrations.

  1. 01

    Label the key

    e.g. cursor-integration

  2. 02

    Create & copy

    nxs_live_… shown once

  3. 03

    Select in chat

    Or paste into an external client

  4. 04

    Follow usage

    Last used · revoke anytime

// SECTION 06 — ORGANIZATION & USERS

One organisation. Clear roles. Groups that stick.

An owner sits at the top. Admins manage seats and limits. Users sign in, join permission groups, and work under the quotas you set — invite links or direct Add User.

A · Control directory, live

Invite, group, and seat every person.

The real User Management screen: invite links, groups, the directory — then New user, and the User Detail drawer (Overview, API keys, Usage). Hover to pause; click a step to jump.

phronexus control · users · directory
Search users…
AO
// ACCESS · CONTROL

USER MANAGEMENT

Invite links set role and limits; users choose their profile at signup. Direct Add User creates an active account immediately.

// INVITE · LINKS

Invite links

Create a reusable signup code for northwind ops. Set role and limits here — the user picks their name and email when they sign up.

No active invites. Create one to share a signup link.

Need to set the password yourself? Use Add User in the table below.

TOTAL USERS
03
PRESENCE
01
BLOCKED REQUESTS
0
// ONLINE · PRESENCE
  • Alex Ortega·just now
// GROUPS

Permission groups

Select users in the table, then add them to a group. Models are allowed for everyone unless you restrict a model on a node.

Group name
  • ops-core
    2 members · LLM on
  • vision-pilot
    1 member · LLM on
// DIRECTORY
USER TABLE
3 / 3 USERS
Search name or email…
All rolesAll statuses
USERROLEACTIONS
AO
Alex Ortega
alex@northwind.ops
Owner
MC
Maya Chen
maya@northwind.ops
Admin
JL
Jordan Lee
jordan@northwind.ops
User
NEXUS CONTROL · RBAC ENFORCED · ALL ACCESS EVENTS LOGGED

B · Organisation structure

Owner → admins → users → groups.

northwind ops is one organisation. The owner owns the tenant. Admins manage seats, invites, and limits. Users sign in under those limits. Permission groups (ops-core, vision-pilot) batch LLM access so you do not edit every seat by hand.

LayerWhat it does
OrganisationTenant boundary · name · invite codes
OwnerFull control · cannot be demoted
AdminUsers, groups, limits, invites
User / RestrictedPortal access under quota
Permission groupsBatch LLM access · assign from the table
User detailOverview · API keys · usage log per seat

C · How seats get created

Invite link — or Add User with a password.

Invite links bake role and limits into a signup code; the person chooses name and email. Add User creates an active account immediately when you need to set the password yourself.

  1. 01

    Create invite or Add User

    Role · token · prompt · rate limits

  2. 02

    They sign in

    Portal chat under those quotas

  3. 03

    Optional: group them

    Select rows → Add to ops-core

  4. 04

    Revoke or suspend anytime

    Directory actions · audit logged

// SECTION 07 — MONITORING & ANALYTICS

Your seat — or the whole organisation.

Portal Usage is personal: quota, recent requests, and drill-down by API key. Control Analytics is org-wide: tokens, cost, top users, and traffic across every seat.

A · Two scopes

Your seat — or the whole organisation.

Portal Usage is personal. Control Analytics is organisation-wide. They never mix in the same screen — each person sees their own quota and keys; admins see the fleet.

SurfaceAudienceShows
Portal · UsageEach userQuota · requests · usage by key
Control · AnalyticsOwner / adminTokens · cost · top users · traffic

B · Per user (Portal)

Your usage & groups · requests · by key.

Each person sees their own monthly quota, permission groups, a filterable request history, and charts scoped to a single API key — without anyone else's traffic.

Portal · Your usage
// PORTAL · PERSONAL
Portal · Your usage

Usage

Quota, charts, and request history

AO
alex@northwind.ops
Owner
ACCOUNT

Your usage & groups

Monthly quota across all API keys, permission groups, and activity over the last 30 days.

MONTHLY QUOTA
User tokens12K / 100K tok
60 req/min8,192 char prompt max
PERMISSION GROUPS

No permission groups assigned

Groups control model access and LLM permissions. Contact an admin to change membership.

REQUESTS (30D)
39
Last active 12h ago
TOKENS (30D)
12K
100% success rate
SUCCESS
39
BLOCKED / ERROR
0 / 0
DAILY TOKENS · 30 DAYS

C · Organisation (Control)

Fleet-wide usage analytics.

Admins open Usage Analytics for the rolling window: total tokens, estimated cost, request volume, model mix, and a top-users table.

Control · Usage analytics
// CONTROL · ORGANISATION
Control · Usage analytics
// INFERENCE · USAGE

USAGE ANALYTICS

Live token consumption · model breakdown · top users · rolling window · UTC

7d30d
WINDOW · 7d · COST ESTIMATED FROM TOKENS
19 Aug → 25 Aug 2026 UTC
TOTAL TOKENS
12.3K
▲ 100%
EST. COST
$0.02
▲ 100%
REQUESTS
39
▲ 100%
AVG TOKENS / REQ
315
39 ok · 0 err
CHART.01
TOKEN USAGE
7-day trend
CHART.02
TOKENS BY MODEL
7d
CHART.03
DAILY REQUESTS
7d
SPLIT
MODEL SHARE
by tokens
Qwen2.5-VL-7B…48.3%
Llama-3.1-8B…27%
Coder-14B…15%
R1-Distill…9.7%
// CONSUMERS
TOP USERS
USERTOKENSCOST
Alex Ortega4.8K$0.01
Maya Chen3.1K$0.01
Jordan Lee2.4K<$0.01
// TRAFFIC
REQUEST TYPES
chat 39·12.3K tok

Cost figures are estimated from token volume until per-model billing rates are configured.

NEXUS CONTROL · LIVE USAGE EVENTS · UTC TIMESTAMPS
// SECTION 08 — INTEGRATIONS

Your existing tools, pointed at your own models.

An OpenAI-compatible API means anything already built for OpenAI's format works here with a one-line change: the URL.

A · OpenAI-compatible API

One Base URL. Your existing tools.

Paste /v1, a Nexus key, and a published model id. Same allowlists and quotas as the portal — Cursor, Open WebUI, Continue, and anything else that speaks OpenAI’s format work without a rewrite.

phronexus portal · openai-compatible api
EXTERNAL APPS

OpenAI-compatible API

Connect Open WebUI, Continue, Cursor, or any OpenAI client to your allowlisted local models through Nexus. Every call uses an API key so usage stays fully traceable.

API ORIGIN (REACHABLE FROM THE CLIENT)
http://10.20.14.8:3001

WSL and other machines cannot use 127.0.0.1 — pick the LAN / vEthernet address.

BASE URL
http://10.20.14.8:3001/v1
API KEY
nxs_live_… (from Portal → API keys)

Authorization: Bearer · or X-Api-Key

OPEN WEBUI / CURSOR / OPENAI CLIENTS
  1. Settings → Connections → OpenAI → add connection
  2. API URL: paste the Base URL above (must end with /v1)
  3. API key: your Nexus key (starts with nxs_live_)
  4. Pick a model from the list below in the client
MODEL IDS (USE IN CLIENT)
Refresh
LocalServer_Qwen2.5-VL-7B-Instruct
Local Server · Qwen2.5-VL-7B-Instruct-Q4_K_M.gguf
Copy id
TEST WITH CURL
Copy
curl http://10.20.14.8:3001/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer YOUR_NXS_KEY" \
  -d '{"model":"LocalServer_Qwen2.5-VL-7B-Instruct","messages":[{"role":"user","content":"Hello"}]}'
GET /v1/models·POST /v1/chat/completions·Legacy alias: /api/inference/v1/*

B · Open source & private clients

From LibreChat to Cursor — same endpoint.

Open-source frontends and private IDEs both connect as OpenAI clients. Nexus does not care which logo is on the window; it cares that the key is valid and the model is allowlisted.

Clients

Same /v1 endpoint — open source and private tools

Point any OpenAI-format client at Nexus. Allowlists and quotas still apply.

Open source

Open WebUI
Open source
Continue
Open source
LibreChat
Open source
LobeChat
Open source
AnythingLLM
Open source
SillyTavern
Open source
Aider
Open source
Open Interpreter
Open source
Open WebUI
Open source
Continue
Open source
LibreChat
Open source
LobeChat
Open source
AnythingLLM
Open source
SillyTavern
Open source
Aider
Open source
Open Interpreter
Open source

Private & commercial

Cursor
Private
VS Code
Private
JetBrains AI
Private
Windsurf
Private
Zed
Private
Cody
Private
Claude Desktop
Private
Raycast AI
Private
Cursor
Private
VS Code
Private
JetBrains AI
Private
Windsurf
Private
Zed
Private
Cody
Private
Claude Desktop
Private
Raycast AI
Private

C · Connect in four steps

Base URL, key, model — done.

Create a key in the portal, copy the Base URL ending in /v1, pick a published model id, and drop them into the client’s OpenAI connection settings.

  1. 01

    Create a key

    Portal → API keys · nxs_live_…

  2. 02

    Copy Base URL

    …/v1 — reachable from the client

  3. 03

    Paste into the app

    OpenAI connection · Bearer auth

  4. 04

    Select the model id

    From the published list on Nexus

Five ways to run Nexus.

Whatever the shape of your business.

See it runningon your own hardware.

Twenty minutes. We look at how your team works, tell you what it would cost, and tell you honestly if you should wait.

NO DECK · NO PITCH · NO FOLLOW-UP SEQUENCE