Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
107 commits
Select commit Hold shift + click to select a range
7299bec
feat(installer): prepare DGX Station host prerequisites
ericksoa Jul 16, 2026
35c26d3
fix(installer): harden station host preparation
ericksoa Jul 16, 2026
9763bff
test(installer): cover station resume symlinks
ericksoa Jul 16, 2026
8e5fb90
test(installer): cover station resume boundaries
ericksoa Jul 16, 2026
34ac3fb
fix(installer): secure station bootstrap state
ericksoa Jul 16, 2026
c56f6c5
fix(installer): recheck workloads before runtime changes
ericksoa Jul 16, 2026
88d1156
test(installer): prove Station helper ships in bootstrap
ericksoa Jul 16, 2026
7c2ee15
ci: retry credentialed e2e after runner loss
ericksoa Jul 16, 2026
80fdb0a
fix(installer): constrain Docker runtime repair
ericksoa Jul 16, 2026
fb2be36
test(installer): guard Station project assignment
ericksoa Jul 16, 2026
fb6adef
fix(installer): complete Station GPU prerequisites
ericksoa Jul 16, 2026
c1edd7b
fix(installer): preserve Station host policies
ericksoa Jul 16, 2026
b2230f9
fix(installer): persist Station CDI lifecycle
ericksoa Jul 16, 2026
26eef09
fix(installer): enforce DGX Station preparation boundaries
senthilr-nv Jul 16, 2026
a560a5d
feat(vllm): add probe-gated dual DGX Station serving
ericksoa Jul 16, 2026
c5e6ab8
fix(installer): tighten Station host safety gates
senthilr-nv Jul 16, 2026
f3c6f13
Merge remote-tracking branch 'upstream/main' into feat/dgx-station-ho…
senthilr-nv Jul 16, 2026
a6d8a18
Merge branch 'main' into feat/dual-dgx-station-inference
ericksoa Jul 16, 2026
541d787
fix(vllm): address dual-station CodeQL findings
ericksoa Jul 16, 2026
6e7a98f
Merge branch 'main' into feat/dual-dgx-station-inference
ericksoa Jul 16, 2026
81c80f3
fix(installer): fail closed on Station CDI refresh
senthilr-nv Jul 16, 2026
e5546b4
Merge branch 'main' into feat/dual-dgx-station-inference
ericksoa Jul 16, 2026
7304a11
merge: refresh Station host preparation onto main
senthilr-nv Jul 16, 2026
a0e7a79
Merge branch 'feat/dgx-station-host-prereqs' into feat/dual-dgx-stati…
ericksoa Jul 16, 2026
01ddd1f
fix(installer): avoid Docker restart on workload race
senthilr-nv Jul 16, 2026
edb7185
Merge branch 'feat/dgx-station-host-prereqs' into feat/dual-dgx-stati…
ericksoa Jul 16, 2026
cf46723
feat(installer): prepare trusted dual DGX Stations
ericksoa Jul 16, 2026
21f906a
test(installer): keep Station fixtures linear
ericksoa Jul 16, 2026
0e36663
Merge remote-tracking branch 'origin/main' into feat/dgx-station-host…
ericksoa Jul 16, 2026
ee5d6d3
fix(installer): clarify Station GPU probe contract
ericksoa Jul 16, 2026
230f0cb
chore(installer): sync DGX Station host prerequisites
ericksoa Jul 16, 2026
826b7b8
fix(inference): pin qualified Station SSH identity
ericksoa Jul 16, 2026
4cfc457
Merge branch 'main' into feat/dual-dgx-station-inference
ericksoa Jul 16, 2026
9e0cdbd
fix(installer): preserve provider help parity
ericksoa Jul 16, 2026
b5bb20f
test(inference): add isolated dual Station simulator
ericksoa Jul 17, 2026
3ac5606
Merge branch 'main' into feat/dual-dgx-station-inference
ericksoa Jul 17, 2026
17015dd
Merge branch 'main' into feat/dual-dgx-station-inference
ericksoa Jul 17, 2026
9e47986
fix(inference): harden dual Station runtime
ericksoa Jul 17, 2026
dc4a893
Merge branch 'main' into feat/dual-dgx-station-inference
ericksoa Jul 17, 2026
c1613c0
Merge branch 'main' into feat/dual-dgx-station-inference
ericksoa Jul 17, 2026
3aa447f
Merge branch 'main' into feat/dual-dgx-station-inference
ericksoa Jul 17, 2026
e447a84
fix(vllm): harden dual Station lifecycle
ericksoa Jul 17, 2026
92d7f2a
Merge origin/main into feat/dual-dgx-station-inference
ericksoa Jul 17, 2026
152ebd9
test(vllm): satisfy conditional growth guardrail
ericksoa Jul 17, 2026
66af9ec
Merge origin/main into feat/dual-dgx-station-inference
ericksoa Jul 17, 2026
2b94049
test(vllm): keep fixture support out of build
ericksoa Jul 17, 2026
30411b8
test(vllm): secure Station lock fixture
ericksoa Jul 17, 2026
6c3fc06
fix(vllm): harden Station preflight validation
ericksoa Jul 17, 2026
8f210a0
test(vllm): fail closed on Station fixtures
ericksoa Jul 17, 2026
8fa029e
fix(vllm): reject conflicting Station peer model
ericksoa Jul 17, 2026
31a2591
Merge origin/main into feat/dual-dgx-station-inference
ericksoa Jul 17, 2026
7bc49df
docs(vllm): correct dual Station network boundary
ericksoa Jul 17, 2026
634df21
Merge origin/main into feat/dual-dgx-station-inference
ericksoa Jul 17, 2026
fb27f56
test(vllm): pin dual Station port isolation
ericksoa Jul 17, 2026
7413003
test(vllm): cover attached Docker publish flags
ericksoa Jul 17, 2026
686f68a
Merge main into feat/dual-dgx-station-inference
ericksoa Jul 17, 2026
be74387
fix(installer): recognize P3830 DGX Station identity
ericksoa Jul 17, 2026
d8fe32e
Merge origin/main into feat/dual-dgx-station-inference
ericksoa Jul 17, 2026
d672641
merge: integrate current main into dual Station inference
ericksoa Jul 20, 2026
61afaf9
docs: describe automatic dual Station selection
ericksoa Jul 20, 2026
c1ab8c9
docs: repin Station prompt asset
ericksoa Jul 20, 2026
1dffcfe
merge: refresh dual Station branch from main
ericksoa Jul 20, 2026
62c1eb4
docs: refresh Station prompt asset pin
ericksoa Jul 20, 2026
e664b19
test: keep ownership fixture linear
ericksoa Jul 20, 2026
f05cec2
fix: scope controller binding dependency
ericksoa Jul 20, 2026
c90f795
test: provide Ultra token in storage fixtures
ericksoa Jul 20, 2026
bc3049d
merge: refresh dual Station branch from main
ericksoa Jul 20, 2026
43105d2
feat(vllm): align dual Station runtime with playbook
ericksoa Jul 21, 2026
f3fa5ae
merge: refresh dual Station branch from main
ericksoa Jul 21, 2026
962f80c
merge: refresh dual Station branch from main
ericksoa Jul 21, 2026
65bc93c
test(docs): refresh merged Station prompt pin
ericksoa Jul 21, 2026
5cd7535
test(vllm): align Station checks with Ray runtime
ericksoa Jul 21, 2026
09d5791
test(vllm): remove dual Station simulator
ericksoa Jul 21, 2026
829855b
test(vllm): satisfy fixture growth guard
ericksoa Jul 21, 2026
39138da
fix(station): tighten preparation prerequisites
ericksoa Jul 21, 2026
8537198
fix(installer): retain incomplete Station pair state
ericksoa Jul 21, 2026
abc34e4
Merge branch 'main' into feat/dual-dgx-station-inference
ericksoa Jul 22, 2026
b97079c
merge(main): refresh dual DGX Station inference
senthilr-nv Jul 25, 2026
6c93876
merge(main): integrate gateway service ownership
senthilr-nv Jul 25, 2026
131129a
fix(vllm): harden managed Station recovery
senthilr-nv Jul 25, 2026
fff4087
docs(station): repin corrected prompt asset
senthilr-nv Jul 25, 2026
94a2ea4
merge(main): refresh dual Station inference
senthilr-nv Jul 25, 2026
bf46e62
docs(station): clarify Ultra topology selection
senthilr-nv Jul 25, 2026
976112b
docs(station): repin Ultra topology prompt
senthilr-nv Jul 25, 2026
84befb3
merge(main): refresh dual Station inference
senthilr-nv Jul 25, 2026
60ecde1
fix(vllm): complete fresh dual Station preparation
senthilr-nv Jul 26, 2026
88c483e
merge(main): refresh dual Station inference
senthilr-nv Jul 26, 2026
4af77a4
fix(station): preserve dual-node resume validation
senthilr-nv Jul 26, 2026
0d62d2c
merge(main): refresh dual Station inference
senthilr-nv Jul 26, 2026
a8f3515
fix(station): normalize local vllm resume path
senthilr-nv Jul 26, 2026
546e67b
fix(station): validate unfiltered peer neighbors
senthilr-nv Jul 26, 2026
93561d8
merge(main): refresh dual Station inference
senthilr-nv Jul 26, 2026
7bf5710
merge(main): refresh dual Station inference
senthilr-nv Jul 26, 2026
39a2a00
merge(main): refresh dual Station inference
senthilr-nv Jul 26, 2026
9d965a2
merge(main): refresh dual Station inference
senthilr-nv Jul 27, 2026
f9bd257
merge(main): refresh dual Station inference
senthilr-nv Jul 27, 2026
e3e58be
merge(main): refresh dual Station inference
senthilr-nv Jul 27, 2026
4485c54
merge(main): refresh dual Station inference
senthilr-nv Jul 27, 2026
950cb3a
fix(vllm): keep Nemotron Ultra public
senthilr-nv Jul 27, 2026
c8998d6
fix(onboard): preserve dual Station model intent
senthilr-nv Jul 27, 2026
432ab5e
fix(vllm): accept local default Docker context
senthilr-nv Jul 27, 2026
b46549e
fix(vllm): accept detected local Docker socket
senthilr-nv Jul 27, 2026
b7ffe7d
fix(vllm): probe Station GPU runtime directly
senthilr-nv Jul 27, 2026
d9008b6
fix(vllm): accept Station route JSON variants
senthilr-nv Jul 27, 2026
eada676
fix(vllm): accept filtered Station neighbor JSON
senthilr-nv Jul 27, 2026
204798b
fix(vllm): allow Ray on Station runtime tmpfs
senthilr-nv Jul 27, 2026
dd03ed8
fix(vllm): keep FlashInfer cubin links writable
senthilr-nv Jul 27, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
35 changes: 35 additions & 0 deletions docs/get-started/dgx-station-preparation.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -162,6 +162,41 @@ The AI Developer Tools path can also leave the packaged CDI refresh units enable

For the current support status and direct GPU policy boundaries, see [Platform Support](../../reference/platform-support).

## Prepare a Two-Station Pair

Complete NVIDIA's [two-Station CX8 fabric playbook](https://build.nvidia.com/station/connect-two-stations/instructions) before running NemoClaw pair preparation.
The published [dual-Station Nemotron Ultra playbook](https://build.nvidia.com/station/nemoclaw/dual-nodes) is the runtime recipe tracked by this managed path.

When no peer or model is selected explicitly, Station Express selects `nemotron-3-ultra-550b-a55b` and checks for one trusted reciprocal peer at the deterministic counterpart address on each of two configured private `/30` ConnectX-8 rails.
If both Stations pass reciprocal identity, GPU, route, neighbor, and jumbo-frame checks, NemoClaw uses the distributed Ultra recipe.
If no trusted pair qualifies, Express retains the existing single-Station Ultra recipe.
An explicit `NEMOCLAW_VLLM_MODEL` remains authoritative.
When `NEMOCLAW_DGX_STATION_PEER` is set, that exact peer must qualify or setup stops instead of falling back.

Configure both Stations before installation with exactly two active 400 Gbit/s Ethernet rails, MTU 9000, and one usable private `/30` address per rail.
The reciprocal addresses must use direct-link routes, the expected peer MAC neighbors, and jumbo-frame connectivity in both directions.
The SSH target must already have usable host-key trust and non-interactive authentication, and the selected peer account must have passwordless `sudo` for remote preparation.
NemoClaw checks only the two deterministic `/30` counterpart addresses; it does not scan other addresses, configure the rails, enroll SSH trust, or accept a shared `/24` as equivalent evidence.

Station preparation binds the preparing non-root account's UID in root-owned `/etc/nemoclaw/dual-station-controller-uid`.
Rebinding requires an administrator to remove that file before preparation is rerun as the replacement account.
If either Station requires a reboot, the installer stops with status `10` and names the host to reboot manually.
The initial local-host receipt preserves the accepted NemoClaw revision and Express selections.
After reciprocal peer qualification begins, owner-only pair state additionally binds the preparation helper, SSH host key, GPU identities, and reciprocal rails.
The installer never reboots either host automatically.

<Warning title="Trusted Pair Boundary">
The dual-Station runtime uses unauthenticated Ray, NCCL, and vLLM coordination traffic, including the Ray head on TCP port `6379` and Ray's worker traffic.
Treat both Stations and every host that can reach either private rail as mutually trusted.
The operator owns physical rail isolation, host firewalling, SSH trust, and reboot control.
Do not use the pair path on a shared or routed network without separately reviewed isolation evidence.
</Warning>

<Warning title="DGX Station Support Status">
The dual-Station path is a Deferred evaluation.
It does not change the single-Station support status.
</Warning>

## Next Step

After your host passes preparation or validation, continue with the [Quickstart](../quickstart).
35 changes: 29 additions & 6 deletions docs/inference/set-up-vllm.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -19,10 +19,11 @@ The vLLM `/v1/responses` endpoint does not run the configured tool-call parser,

<Warning>
Local vLLM does not require authentication by default.
NemoClaw needs port `8000` on host loopback for validation and on the OpenShell Docker bridge for sandbox traffic.
Existing-server and single-host managed-vLLM paths need port `8000` on host loopback for validation and on the OpenShell Docker bridge for sandbox traffic.
Use a host firewall with default-deny inbound rules.
Allow TCP port `8000` only from the OpenShell Docker subnet to its gateway address, keep loopback access, and deny the port on every other interface.
Do not expose the port to your LAN or the internet.
The qualified dual-DGX Station path instead binds its bearer-protected `/v1` API to one selected private rail address and uses host networking on both runtime containers; follow the additional isolation guidance in the DGX Station section below.
</Warning>

## Use an Existing Server
Expand Down Expand Up @@ -87,6 +88,7 @@ Managed profiles and model-specific recipes use immutable image digests:
- DGX Spark and DGX Station models without a model-specific runtime use the `linux/arm64` digest `sha256:9204569b17ee4c0eff75194b8e6e458479c8aee18953b5ab9cf359fcdac659e2` with a compressed layer size of `9.60 GB` under `nvcr.io/nvidia/vllm:26.05.post1-py3`.
- The DGX Station Nemotron 3 Ultra express recipe uses the multi-platform index digest `sha256:0fec7ec5f3e6bc168e54899935fb0557da908a4832a1dbc88e2debcf2f889416` under `vllm/vllm-openai:v0.22.0`; on DGX Station, that index selects a `linux/arm64` manifest with a compressed layer size of `10.67 GB`.
It also pins Hugging Face revision `183968f87ae4cedce3039313cac1fd43d112c578` for the approximately `352.38 GB` model download.
- A qualified two-Station Ultra pair instead uses the ARM64 manifest digest `sha256:2cc49b81319f7a66a33dd8bd63a7bfddae079122b33ce51989b6828a1f038c37` under `vllm/vllm-openai:v0.25.1-aarch64`, with `10.24 GB` of compressed layers.
- Generic Linux `arm64` hosts use `sha256:447995cbb57e6c7cf792cab95e9852e5f62b5fb6d2f39e030fa4eda9a54eadb4` with a compressed layer size of `9.28 GB` under `nvcr.io/nvidia/vllm:26.03.post1-py3`.
- Generic Linux `amd64` hosts use `sha256:7be6c2f676c36059a494fe17254e69ae5c677535ba6191044e5fc8e42a91c773` with a compressed layer size of `8.93 GB` under `nvcr.io/nvidia/vllm:26.03.post1-py3`.

Expand Down Expand Up @@ -181,7 +183,13 @@ When you start managed vLLM outside the installer express flow, NemoClaw uses th
| DGX Station | `deepseek-ai/DeepSeek-V4-Flash` |
| Linux with an NVIDIA GPU | `nvidia/NVIDIA-Nemotron-3-Nano-4B-FP8` |

On DGX Station, accepting the installer express prompt sets `NEMOCLAW_VLLM_MODEL=nemotron-3-ultra-550b-a55b` and overrides the profile default.
On DGX Station, accepting the installer express prompt selects `NEMOCLAW_VLLM_MODEL=nemotron-3-ultra-550b-a55b` and checks for a trusted reciprocal pair.
The installer prepares the local host with the pinned Station helper, derives exactly one counterpart from each of two configured private `/30` CX-8 rails, and consults existing SSH trust only for those two addresses.
At least one derived address must already be trusted; if both are trusted, their host keys must identify one coherent SSH host.
After the Stations pass reciprocal GPU, rail, MAC, route, neighbor, and jumbo-frame checks across both rails, the installer prepares the peer with the same reviewed helper and exports the qualified peer.
No managed-vLLM runtime image or model download starts before the pair-qualification or single-Station fallback decision completes.
If no trusted pair qualifies, Express retains the existing single-Station Ultra recipe.
Setting `NEMOCLAW_VLLM_MODEL` remains authoritative; setting `NEMOCLAW_DGX_STATION_PEER` requests that exact already-trusted peer and fails closed if it does not qualify.
To select the existing `deepseek-v4-flash` recipe while retaining the same one-confirmation express flow, run:

```bash
Expand All @@ -200,9 +208,22 @@ curl -fsSL https://www.nvidia.com/nemoclaw.sh | \

The flag requires an interactive terminal; in a `curl | bash` pipeline, `/dev/tty` must be available.
Without terminal access, the installer stops before it installs Docker or build dependencies instead of silently continuing with another configuration.
The registered Ultra recipe tracks the [official DGX Station deployment guide](https://github.com/NVIDIA-NeMo/Nemotron/blob/287ae845639d2ce998998cb8fd1f70a3fa943c0b/usage-cookbook/Nemotron-3-Ultra/StationDeploymentGuide/README.md) and configures the pinned model revision, CPU offload, `16 GB` of shared memory, memory/stack ulimits, MTP speculative decoding, and the Nemotron reasoning and tool-call parsers.
NemoClaw intentionally keeps its existing bridge-networked managed-inference topology instead of importing the playbook's host-network setting.
The container publishes port `8000` through Docker, so apply the firewall guidance at the top of this page.
The registered single-Station Ultra recipe tracks the [official DGX Station deployment guide](https://github.com/NVIDIA-NeMo/Nemotron/blob/287ae845639d2ce998998cb8fd1f70a3fa943c0b/usage-cookbook/Nemotron-3-Ultra/StationDeploymentGuide/README.md) and configures the pinned model revision, CPU offload, `16 GB` of shared memory, memory/stack ulimits, MTP speculative decoding, and the Nemotron reasoning and tool-call parsers.
Configure the physical rails and SSH host-key and authentication trust before installation.
NemoClaw does not modify rail configuration, enroll SSH trust, or reboot either host automatically.
If the local host requires a reboot during initial preparation, the installer stops with status `10`; its owner-only receipt preserves the printed exact revision and Express selections.
After reciprocal peer qualification begins, the pair state additionally binds SSH, GPU, and rail identity and names either host that must be rebooted manually.
The registered two-Station Ultra recipe tracks the [latest NVIDIA dual-Station playbook](https://build.nvidia.com/station/nemoclaw/dual-nodes).
It pins vLLM `0.25.1` and Ray `2.56.0`, uses one tensor-parallel rank per Station with pipeline parallelism across the pair, serves the `nemotron-ultra` alias with a `262144`-token model limit, and retains the Nemotron reasoning and tool-call parsers.
The single-Station fallback keeps NemoClaw's bridge-networked managed-inference topology and publishes port `8000` through Docker.
The qualified dual-Station runtime intentionally uses Docker host networking on both containers so NCCL and RDMA can bind the validated direct-attach rails.
The head binds only the selected rank-0 address on a qualified private `/30` rail and requires the generated bearer key for `/v1`; `/health` remains unauthenticated for readiness.
The worker joins the Ray cluster and exposes no vLLM API.
Neither dual-Station container publishes a Docker port, so Docker bridge isolation and port-mapping rules do not protect this path.
As compensating controls, both containers run as the probed non-root UID and GID with a read-only root filesystem, all Linux capabilities dropped, `no-new-privileges`, only the selected GPU UUID and exact `uverbs` devices, and a read-only model cache; the worker does not receive the serving key.
Treat both Stations and their direct rails as one trusted runtime boundary.
Allow port `8000` from the OpenShell Docker subnet only to the selected rank-0 rail address, deny it on management and LAN interfaces, and restrict distributed-serving traffic to the reciprocal addresses on the two qualified private rails.
That distributed traffic includes the Ray head on TCP `6379` and Ray worker ports; do not expose either rail to an untrusted or routed network.

<Warning>
Before managed vLLM setup on DGX Station, follow [Prepare DGX Station to Install NemoClaw](../../get-started/additional-setup/dgx-station-preparation).
Expand Down Expand Up @@ -234,7 +255,9 @@ NEMOCLAW_PROVIDER=install-vllm \
On DGX Spark and DGX Station, `NEMOCLAW_PROVIDER=install-vllm` is sufficient for a non-interactive run.
Add `NEMOCLAW_EXPERIMENTAL=1` on a generic Linux NVIDIA GPU host.
Non-interactive runs use the profile default unless you set `NEMOCLAW_VLLM_MODEL`.
On DGX Station, a direct provider-only run therefore selects `deepseek-v4-flash`; the installer express flow sets the Nemotron 3 Ultra override for you.
The commands above invoke `nemoclaw onboard` directly, so a DGX Station run with no model or peer selects the `deepseek-v4-flash` profile default.
Supplying `NEMOCLAW_PROVIDER=install-vllm` to the shell installer enters the Station host-preparation boundary while retaining that profile default.
To request a non-interactive pair through this path, also set `NEMOCLAW_DGX_STATION_PEER`; the exact peer must qualify, and the installer selects the distributed Nemotron 3 Ultra recipe.

For a headless DGX Station setup that selects DeepSeek V4 Flash explicitly, use the environment-variable path instead of `--station-deepseek`.

Expand Down
4 changes: 3 additions & 1 deletion docs/reference/commands.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -3562,7 +3562,9 @@ Set them before running `$$nemoclaw onboard`.
| `SANDBOX_NAME` | sandbox name | Compatibility spelling used after `NEMOCLAW_SANDBOX_NAME` and `NEMOCLAW_SANDBOX`. |
| `NEMOCLAW_INSTALL_REF` | git ref | For internal installer commands: the git ref to install from. A nonempty value takes precedence over `NEMOCLAW_INSTALL_TAG`. Overridden by the `--install-ref` flag. |
| `NEMOCLAW_INSTALL_TAG` | release tag | For internal installer commands: the release tag to install when `NEMOCLAW_INSTALL_REF` is unset or empty. Defaults to the admin-promoted `lkg` tag when unset. Overridden by the `--install-tag` flag. |
| `NEMOCLAW_VLLM_MODEL` | registry slug or Hugging Face model id | Selects the model the managed-vLLM install path serves. Recognised slugs: `qwen3.6-27b`, `qwen3.6-35b-a3b-nvfp4`, `nemotron-3-nano-4b`, `deepseek-v4-flash`, `nemotron-3-ultra-550b-a55b`, `deepseek-r1-distill-70b`. Unset uses the per-platform profile default. The DGX Station express installer sets `nemotron-3-ultra-550b-a55b` explicitly. Gated models (e.g. `deepseek-r1-distill-70b`) require `HF_TOKEN` or `HUGGING_FACE_HUB_TOKEN`. |
| `NEMOCLAW_VLLM_MODEL` | registry slug or Hugging Face model id | Selects the model the managed-vLLM install path serves and remains authoritative during DGX Station installer setup. Recognised slugs: `qwen3.6-27b`, `qwen3.6-35b-a3b-nvfp4`, `nemotron-3-nano-4b`, `deepseek-v4-flash`, `nemotron-3-ultra-550b-a55b`, `deepseek-r1-distill-70b`. Station Express selects `nemotron-3-ultra-550b-a55b`; a qualified reciprocal pair uses the distributed topology, while no qualifying pair retains the single-Station Ultra topology. Outside Station Express, unset uses the per-platform profile default. Gated models (e.g. `deepseek-r1-distill-70b`) require `HF_TOKEN` or `HUGGING_FACE_HUB_TOKEN`. |
| `NEMOCLAW_DGX_STATION_PEER` | SSH host or `user@host` | Selects one exact, already-trusted DGX Station peer for Nemotron 3 Ultra pair qualification. The peer must match the reciprocal private `/30` rail and hardware checks; an explicit peer failure stops setup instead of falling back. NemoClaw does not enroll SSH trust or accept a port or SSH option in this value. When unset, DGX Station installer discovery checks only the two deterministic `/30` counterpart addresses. A peer cannot be combined with an explicit non-Ultra model; conflicting explicit selections fail before pair preparation. |
| `NEMOCLAW_DGX_STATION_SSH_BINDING` | opaque installer-managed token | Carries the qualified peer endpoint and host-key binding from DGX Station pair preparation into the current managed-vLLM install. The installer creates and clears this token; operators should not set or persist it. Missing, changed, or mismatched binding state fails before peer SSH or Docker work. |
| `NEMOCLAW_VLLM_EXTRA_ARGS_JSON` | JSON array of non-blank strings | Appends advanced operator-owned tokens to the managed `vllm serve` command after NemoClaw's registry defaults. Example: `["--max-num-seqs","2"]`. Malformed JSON, non-string tokens, or blank tokens fail before Docker work starts. |
<AgentOnly variant="openclaw">
| `NEMOCLAW_MINIMAL_BOOTSTRAP` | `1` to enable | Skips default OpenClaw workspace-template seeding for new pristine workspaces. Existing files are not deleted; refer to [Understand Runtime Changes](../manage-sandboxes/configure-sandboxes/understand-runtime-changes). |
Expand Down
Loading
Loading