VEP  /  Compute  /  Sovereign on NVIDIA DGX Spark
Reference architecture · on your own hardware

Your AI worker, on one computer in your own office

  • A VEP AI worker is three parts: a memory, a model that thinks, and the software around them.
  • You choose how many of the three sit on your machine: the memory alone, the memory and the model, or all of it.
  • On an NVIDIA DGX Spark, a computer that sits on a desk, all three fit in one box inside your network.

Sovereign here means one thing: the data and the machine are under your control and your law. It is not a certification, and we do not claim one.

Three NVIDIA DGX Spark units stacked on top of each other
NVIDIA DGX Spark. Shown: three units stacked. One unit sits on a desk and runs from a normal wall socket; add units as you grow. NVIDIA and DGX Spark are trademarks of NVIDIA Corporation.

Four words, and the rest of this page makes sense

Watch one AI worker come apart into its three parts. Every section below uses only these four words.

A worker is three partsand you choose where each one lives
Animated diagram: one box labelled AI worker splits into three slabs, platform, intelligence and brain, and joins back together. ONE AI WORKER PLATFORM INTELLIGENCE BRAIN the software around it the app, the tools, the permissions, the log the model that thinks reads, reasons, writes the answer the memory facts it learned, each with its source three parts you can place separately: on your machine, or in VEP's cloud

Brain

The worker's memory: facts it learned, each with its source.

Think of a filing cabinet where every paper is dated and initialled. It stores and finds. It does not think. Change the model tomorrow and the cabinet is untouched.

Intelligence

The thinking part: it reads, reasons, and writes the answer.

Think of a specialist who has read everything in the world and not one page about your company, until the cabinet hands over the papers.

Platform

Tools, the app, the permissions, and a log of everything.

The desk, the phone, the door pass and the visitor book. It decides what the worker may do and writes down what it did.

Closed box

Your own machine, with every outward door shut by default.

Doors open one at a time, because you opened them. Closed means the doors, not the lid: you can inspect everything inside.

It is software, not a person. It only does what your rules allow, and every action is written down.

What happens when someone asks it a question

Watch one request travel through the machine. Every hop stays behind the dotted line; the red doors stay shut unless you open one on purpose.

Inside your network · NVIDIA DGX SparkDiagram · placement 3, everything on your side
Animated diagram: a request goes from your people to the brain, then to the model, then to the runtime, and the answer comes back, all inside the perimeter; outbound doors are closed. YOUR NETWORK · PEOPLE, BRAIN, MODEL, RUNTIME Your people chat · docs · meetings Brain the memory Model thinks on the GB10 GPU Runtime tools · rules · log outbound: shut answer, notes and the log come back to your people over the same internal wire Animated diagram, vertical: a request goes from your people down to the brain, the model and the runtime, and the answer returns; outbound doors are closed. YOUR NETWORK Your people chat · docs · meetings Brain the memory Model thinks on the GB10 GPU Runtime tools · rules · log the answer comes back over the same internal wire outbound doors: shut

1 · Ask

A person writes in the VEP web app on your network, or a document is dropped in.

2 · Remember

The brain finds the facts that matter, with their sources, in the local database.

3 · Think

The model on the GB10 reasons over the question and the facts. No cloud call.

4 · Act and log

The runtime runs the allowed tools, writes new facts back and logs everything for replay.

In kitchen words: a contract arrives and lands on your desk.

This picture shows placement 3, where everything is on your side. In placements 1 and 2 some of these boxes sit in VEP's cloud; the next section shows which.

How much of it sits on your machine?

A worker is three parts: the brain (its memory), the intelligence (the model that thinks) and the platform (the app, tools, permissions and log). Picture a house with three rooms. Green rooms are yours; dashed blue rooms are rented in VEP's cloud.

Three houses, three sets of roomsthe lit house is the one being described
House 1, brain only: the brain room is yours; the intelligence and platform rooms are in VEP's cloud. VEP cloud PLATFORMcloud INTELLIGENCEcloud BRAINyours 1 · Brain only · delivered as a project

Your files and facts stay in your own database. Each question still travels to a model.

House 2, brain plus intelligence: brain and intelligence rooms are yours; the platform room is in VEP's cloud, which only coordinates. VEP cloud PLATFORMcloud INTELLIGENCEyours BRAINyours 2 · Brain + intelligence · beta

The model runs on your GPU, so no question reaches an outside provider. The app and the log stay in VEP's cloud.

House 3, everything on your side: all three rooms are yours and the cloud is faded out; a dotted perimeter surrounds the house. not needed PLATFORMyours INTELLIGENCEyours BRAINyours 3 · Everything on your side · delivered as a project

Memory, model and platform on your machine, behind your firewall. The cloud is not needed.

Which one is for you? Three questions.

Answer about your own company, not about the technology. The matching option lights up below.

1. Must your company's files and facts stay on equipment you own?
2. Are you forbidden to send the question itself to an outside AI provider such as Claude or ChatGPT?
3. Must the chat, the apps and the audit log also sit on your side, with your network closed?
Answer the three questions, or read all three options below.

A starting point, not a quote. We confirm the fit when we scope your data sources.

Delivered as a project

1 · Brain only

For teams whose concern is the memory itself: contracts, files, client facts must never sit in a cloud database, while a cloud model is acceptable for the reasoning.

Brainyour resource (database on your hardware) IntelligenceVEP cloud or your chosen provider (Claude, ChatGPT) PlatformVEP cloud
  • Your facts stay in your database. For each question, only the lines that matter are sent to the model provider together with the question itself. VEP does not keep them outside; what the provider retains is governed by your contract with them.
  • Switch models freely; the brain is untouched.
Delivered as a project

3 · Everything on your side

For closed networks and regulated environments: the whole platform, the brain, the model and the runtime on the Spark, behind your firewall.

Brainyour resource Intelligenceyour GPU Platformyour resource (full VEP installation)
  • Nothing leaves the machine unless you open a door for it: an outbound allow-list is the only way out, and updates arrive as signed bundles.
  • Installed and supported by VEP engineers; a pilot takes weeks.
Beta
It exists and runs today, with our engineers alongside; it is not self-service yet. Expect rough edges; we fix them with you.
Delivered as a project
Our engineers build and install it for you. A pilot is measured in weeks, not an afternoon.
In the standard cloud offer
Available now: workers, brain, authority rules and audit in VEP's cloud.

stays with you VEP cloud

Two things people ask first

Am I locked in? And what can get out? Both answered without a single technical word.

Change the model, keep the memoryany placement
Animated diagram: the brain card with your facts stays lit while the model card is taken out and a different model card is put in its place. BRAIN your facts, with sources your memory stays the model slot MODEL A an open model, today MODEL B a newer one, next year the model is a choice you can change

Whoever sits at the desk, the filing cabinet does not change. Everything the worker learned stays yours.

Two doors: the one you opened, the ones you did notplacement 3
Animated diagram: inside a dotted room, one document bounces off a shut red door and stays inside; another document passes through the one green door you opened. YOUR NETWORK stays inside your mail serveryou opened this one everything elseshut, no key exists closed by default: every open door is a decision you made and can undo

Placement 3, the full installation. In placements 1 and 2 the specialist works outside your building by design.

What is actually on the machine

Everything a VEP worker needs is a set of containers plus a GPU. On a DGX Spark all of it lives on one desk-sized machine inside your network.

The brain

Semantic DNA: every fact the worker knows, with its source and history, in a local PostgreSQL with pgvector. Embeddings are computed locally too.

The model

An open model served on the GB10 GPU from the Spark's 128 GB of unified memory: 30B-class models at 16-bit, 70B-class at 8-bit, up to about 200B with 4-bit quantisation. 128 GB is a hard ceiling; a 70B model at 16-bit needs roughly 140 GB and does not fit. Two linked Sparks reach about 405B, per NVIDIA's published specifications.

The runtime

The worker's reasoning loop, tools and sandbox run next to the model. Chat, documents, meetings and routines are processed on the box.

The workspace

Files, meeting notes, generated documents and the audit trail stay in local object storage; the web app is served from the same machine.

The only doors are the ones you open: an allow-list for outbound connections (your mail server, your CRM, your calendar), an optional VPN for your own people, and the update bundles you fetch and apply yourself. Things every computer needs, such as DNS and time sync, can be pointed at your own internal servers. Remote access for our engineers exists only when you open it. Closed by default. Cloud VEP is never required for the worker to think or remember.

What you get by keeping it in-house

Your data stays yours

Contracts, financials, personnel records: the worker reads them, learns from them and answers about them without any of it crossing your border. You can inspect the design, read the audit log and replay what the worker did. That is not a certification, and we do not claim one.

Your memory stays yours

The memory is the asset. On a Spark it is a database you own, back up, inspect and can export. Switch the model tomorrow and the worker keeps everything it learned.

Predictable cost

One machine, one price, no per-message bill. Whether it is cheaper than a cloud model depends on your volume: we work it out with your numbers during scoping, not with ours on a landing page.

Latency

Model, memory and tools share one bus: no round trip to a cloud model and no provider rate limit. Availability becomes your responsibility; see what you give up.

Same worker, same rules

Delegated authority, human gates, the audit trail with replay, memory rules: the whole VEP method works unchanged. Your people keep the same chat, apps and meetings assistant.

Grows with you

Start with one Spark for a team; link two for larger models; move to a GPU server or a Comino rack later. The installation, the brain and the workers move with you.

What you give up

  • No failover. One Spark is one point of failure; if it stops, the workers stop until it is fixed.
  • Your power, your network. An office outage is now an AI outage.
  • Repairs run on your clock. Someone holds a spare, or accepts the wait for a replacement.
  • Backups are yours. The brain is a database you must back up, and practise restoring.

You are not buying more uptime. You are swapping a provider's rare outage for one you control and must staff.

The computer itself

NVIDIA's desktop AI computer, built for exactly this class of workload. Figures below are NVIDIA's published specifications, not VEP measurements.

ComputeGB10 Grace Blackwell Superchip, up to 1 petaFLOP of FP4 AI performance (NVIDIA figure)
Memory128 GB unified LPDDR5x shared by CPU and GPU: the whole model plus the worker's context in one memory space
ModelsUp to about 200B parameters on one unit with 4-bit quantisation; two units linked over ConnectX networking reach about 405B (NVIDIA figures, quantised)
StorageUp to 4 TB NVMe for models, memory database and workspace
Form factorSits on a desk, runs from a standard power outlet, DGX OS (Ubuntu-based) with the NVIDIA AI stack preinstalled
Fit for VEPA single box runs the platform containers, a 30B-class model at 16-bit or a 70B-class model at 8-bit with room for context, local embeddings and speech, and several workers side by side

How we set it up, step by step

A sovereign installation is an engineering project we run with you, not a download. This is the sequence.

1

Scope

Which teams, which workers, which data sources, which outbound doors (if any), which of the three placements. We size the model to the work.

2

Install

The VEP platform containers go onto the Spark (DGX OS). The brain database, object storage and the web app are initialised inside your network. No cloud account is required for placement 3.

3

Load the model

We pull the chosen open model onto the GB10 and bind the workers' runtime to it. Embeddings and, if wanted, speech recognition run locally as well.

4

Teach the workers

Documents, mail archives, task systems and meeting notes are ingested into Semantic DNA on the box, with the same rules as in the cloud: facts with sources, dedup and supersede chains, human corrections.

5

Open doors deliberately

Each integration (your mail server, CRM, calendar) is an explicit outbound allow-list entry. Everything else stays closed. Updates arrive as signed bundles you apply yourself.

6

Operate

Your people use the same VEP web app, chat, apps and meetings assistant. Audit, replay and authority rules are enforced on the box. We support it under your terms of access.

The questions everyone asks

Can we keep only our files and facts in-house?

Yes. This is placement 1, brain only: the worker's memory (Semantic DNA, facts with their sources) lives in a database on your own hardware while the model runs in VEP's cloud or with a provider you choose. For each question, only the lines that matter travel to the model together with the question itself. Delivered as a project.

Can the thinking part run here too?

Yes. This is placement 2, brain plus intelligence: a GPU machine such as NVIDIA DGX Spark serves an open model locally through the VEP node agent, so no question reaches an outside model provider. VEP's cloud still runs the web app, the apps and the audit log. This lane is in beta: it exists and runs with our engineers alongside, not as self-service.

Can it work with no internet at all?

In placement 3, everything on your side, the whole platform, the brain, the model and the runtime are installed on your DGX Spark or GPU server inside your perimeter. Outbound connections are an allow-list you write; updates arrive as signed bundles you apply yourself; things every computer needs, such as DNS and time sync, can be pointed at your own servers. Delivered as a project; a pilot takes weeks.

How big a model fits on one of these?

One DGX Spark has 128 GB of unified memory. That fits 30B-class open models at 16-bit, 70B-class at 8-bit and up to about 200B parameters with 4-bit quantisation. 128 GB is a hard ceiling: a 70B model at 16-bit needs roughly 140 GB and does not fit. Two linked units reach about 405B, per NVIDIA's published specifications.

Does anything leave the machine?

It depends which of the three placements you choose. In placement 3, everything on your side, nothing by default: prompts, documents, facts and answers stay on the machine, only the integrations you explicitly allow (your mail server, CRM, calendar) get an outbound door, and updates arrive as signed bundles you apply yourself. In placement 1, brain only, each question and the facts it needs go to a model provider you choose. In placement 2, brain plus intelligence, the question stays on your GPU while VEP's cloud runs the app and the audit log.

What happens if the machine fails?

One DGX Spark is one point of failure. If it stops, the workers on it stop until it is repaired or replaced; there is no automatic failover in a single-box installation. Your power and your network become the worker's power and network, repairs run on your clock, and the brain is a database you must back up and practise restoring. You are swapping a provider's rare outage for one you control and must staff.

Is it a worse assistant than the cloud one?

The method is identical: memory with sources, delegated authority, human gates, audit with replay, chat, documents, meetings and App Store apps. What changes is the model behind it: an open model on your GPU instead of a cloud provider such as Claude or ChatGPT. We size the model to your work so the difference is one you choose, not one you discover.

What this is not

  • Not certified against any security standard. We claim no HIPAA, GDPR, SOC 2 or air-gap attestation.
  • It does not make you compliant with anything by itself.
  • The model on your GPU is an open model, not Claude or ChatGPT.
  • In placements 1 and 2, part of the work happens outside your network by design.
  • NVIDIA figures are NVIDIA's published specifications, not VEP measurements.
  • Special-category data such as health records is a decision for your own data protection officer.

Tell us what has to stay in-house

What happens next, in three steps:

Contact sales   All compute options