Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion ci/platform-matrix.json
Original file line number Diff line number Diff line change
Expand Up @@ -129,7 +129,7 @@
"name": "Local NVIDIA NIM",
"status": "experimental",
"endpoint_type": "Local OpenAI-compatible",
"notes": "Requires `NEMOCLAW_EXPERIMENTAL=1` and a NIM-capable NVIDIA GPU. Host must have the NVIDIA Container Toolkit installed and a CDI spec present (`onboard` asserts CDI presence with `assertCdiNvidiaGpuSpecPresent`, `src/lib/onboard/fatal-runtime-preflight.ts`). NIM images pull from `nvcr.io` and require NGC registry login. NemoClaw gates this path behind the experimental flag because it does not auto-select a NIM image for the host today. You must explicitly pick from the validated image list. Managed vLLM has host-specific default models and is not gated on the same boxes. Validated images referenced in `src/lib/inference/config.ts` and `nemoclaw/src/index.ts`: `nvidia/nemotron-3-super-120b-a12b` (default cloud model), `nvidia/nemotron-3-nano-30b-a3b`, `nvidia/llama-3.3-nemotron-super-49b-v1.5`."
"notes": "Requires `NEMOCLAW_EXPERIMENTAL=1` and a NIM-capable NVIDIA GPU. Host must have the NVIDIA Container Toolkit installed and a CDI spec present (`onboard` asserts CDI presence with `assertCdiNvidiaGpuSpecPresent`, `src/lib/onboard/fatal-runtime-preflight.ts`). NIM images pull from `nvcr.io` and require NGC registry login. NemoClaw gates this path behind the experimental flag because it does not auto-select a NIM image for the host today. You must explicitly pick from the validated image list. On Linux arm64 DGX Spark and DGX Station hosts, onboarding warns that some NIM images may not publish a `linux/arm64` manifest; the warning is advisory, and the selected image pull can still fail when the registry has no matching platform manifest. Managed vLLM has host-specific default models and is not gated on the same boxes. Validated images referenced in `src/lib/inference/config.ts` and `nemoclaw/src/index.ts`: `nvidia/nemotron-3-super-120b-a12b` (default cloud model), `nvidia/nemotron-3-nano-30b-a3b`, `nvidia/llama-3.3-nemotron-super-49b-v1.5`."
},
{
"name": "Local vLLM (already running)",
Expand Down
52 changes: 35 additions & 17 deletions docs/about/release-notes.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -16,6 +16,41 @@ NVIDIA NemoClaw is available in early preview starting March 16, 2026.
Use this page to track the highlights of the latest release.
For more detailed release notes, refer to the [NemoClaw GitHub announcements](https://github.com/NVIDIA/NemoClaw/discussions/categories/announcements?discussions_q=is%3Aopen+category%3AAnnouncements).

## v0.0.76

NemoClaw v0.0.76 adds opt-in OTLP observability and a dedicated documentation guide for LangChain Deep Agents Code, makes Nemotron 3 Ultra the managed NVIDIA Endpoints default for Deep Agents Code, contains shared-gateway inference-route conflicts, and improves upgrade recovery, MCP diagnostics, local inference, messaging, and day-two commands.

- LangChain Deep Agents Code now defaults to `nvidia/nemotron-3-ultra-550b-a55b` for NVIDIA Endpoints and uses the native Nemotron 3 Ultra profile from hash-locked Deep Agents Code `0.1.34`.
Managed interactive sessions preserve the optional first-run name prompt while skipping dependency and model selection, and existing sandboxes should be rebuilt after upgrading.
The docs site now publishes a dedicated Deep Agents guide.
For more information, refer to [Quickstart with LangChain Deep Agents Code](/user-guide/deepagents/get-started/quickstart) and [NemoClaw Inference Options](../inference/inference-options).
- Deep Agents sandboxes can opt into backend-neutral OTLP/HTTP tracing during onboarding or rebuild with `--observability`.
The managed exporter sends bounded model and tool content only to a fixed host-local receiver, while an operator-managed collector holds remote backend credentials and controls forwarding.
Treat exported trace content as sensitive application data.
For more information, refer to [Quickstart with LangChain Deep Agents Code](/user-guide/deepagents/get-started/quickstart), [Credential Storage](../security/credential-storage), and [Network Policies](../reference/network-policies).
- Shared OpenShell gateways now reject inference-route changes that conflict with another registered sandbox, including stopped sandboxes.
Compatibility checks cover provider, model, normalized custom endpoint, API family, legacy metadata, and gateway binding before onboarding, `inference set`, or connect-time repair changes state.
`inference set` also accepts onboarding provider aliases and gives actionable model-only and Deep Agents re-onboarding guidance, while OpenClaw onboarding validates Anthropic-compatible streaming event sequences by default.
For more information, refer to [Switch Inference Providers](../inference/switch-inference-providers), [NemoClaw Inference Options](../inference/inference-options), and [Troubleshooting](../reference/troubleshooting).
- Installer upgrades require a fresh backup of every registered sandbox before replacing the gateway.
Pre-fingerprint OpenClaw and Hermes sandboxes can recover only after exact-name managed-image confirmation, recorded custom-image cases remain blocked, and successful recovery restores the recorded provider and route without starting generic onboarding.
For more information, refer to [Manage Sandbox Lifecycle](../manage-sandboxes/lifecycle), [NemoClaw CLI Commands Reference](../reference/commands), and [Credential Storage](../security/credential-storage).
- Startup and onboarding recover more bounded local drift.
OpenClaw can bootstrap the same-device CLI pairing request when device listing is itself pairing-gated, startup safely reclaims an exact root-owned mutable-config posture while rejecting ambiguous states, and interrupted resumable onboarding points to `onboard --resume`.
For more information, refer to [Security Best Practices](../security/best-practices), [Troubleshooting](../reference/troubleshooting), and [NemoClaw CLI Commands Reference](../reference/commands).
- Managed MCP diagnostics can now verify OpenShell credential replacement at the wire boundary.
`mcp add` and single-server `mcp status` run a gated differential probe and report verified or inconclusive credential resolution without capturing endpoint response bodies; `--probe` and `--no-probe` control the check.
For more information, refer to [Set Up MCP Servers](../manage-sandboxes/set-up-mcp-servers) and [NemoClaw CLI Commands Reference](../reference/commands).
- Managed local inference enables automatic tool choice with the `qwen3_coder` parser for the generic-Linux Nemotron 3 Nano vLLM default.
Linux arm64 DGX Spark and DGX Station onboarding warns that some NIM images may lack an arm64 manifest, and troubleshooting now distinguishes direct sandbox DNS limitations from policy-covered inference, messaging, and search health.
For more information, refer to [Use a Local Inference Server](/user-guide/openclaw/inference/use-local-inference), [NemoClaw Inference Options](../inference/inference-options), and [Troubleshooting](../reference/troubleshooting).
- Messaging setup handles more runtime and installation edge cases.
OpenClaw WhatsApp pairing renders a terminal QR code with a four-module quiet zone, Teams plugin installation tolerates verbose npm metadata, and newly generated OpenClaw config no longer references the uninstalled `qqbot` plugin.
For more information, refer to [Messaging Channels](/user-guide/openclaw/manage-sandboxes/messaging-channels).
- Day-two commands behave more predictably in automation and across agents.
Non-terminal `exec` closes stdin unless `--stdin` explicitly forwards a pipe, session passthrough selects the sandbox's native agent binary, resumed Hermes one-shot turns append to the selected session, and `gc` finds orphaned gateway-built and locally prebuilt sandbox images.
For more information, refer to [NemoClaw CLI Commands Reference](../reference/commands) and [Manage Sandbox Lifecycle](../manage-sandboxes/lifecycle).

## v0.0.75

NemoClaw v0.0.75 upgrades the bundled OpenClaw runtime to `2026.6.10` and improves sandbox upgrade and prepared-backup recovery, custom endpoint inference routing, and local Docker-driver sandbox JWT handling.
Expand Down Expand Up @@ -48,23 +83,6 @@ NemoClaw v0.0.74 upgrades the OpenShell policy boundary, adds managed MCP and pr
The selected mode persists through resume and transactional rebuilds, and model-specific compatibility safeguards can keep an incompatible model on direct disclosure.
Sandbox-first `inference get` and `inference set` commands now provide the same route controls as their global forms.
For more information, refer to [Tool Calling Reliability](/user-guide/openclaw/inference/tool-calling-reliability), [Model Capability Audit](../inference/model-capability-audit), and [NemoClaw CLI Commands Reference](../reference/commands).
- Shared OpenShell gateways now enforce a single compatible inference route across every registered sandbox, including stopped sandboxes.
Onboarding, connect-time repair, and `inference set` reject provider/model conflicts; custom routes must also match the normalized endpoint and API family.
As a migration requirement, custom switches must provide `--endpoint-url` and an unambiguous API family, and incomplete legacy custom-route metadata fails closed until the sandbox is removed and re-onboarded with complete metadata.
After backing up an affected workspace, an OpenAI-compatible route can be re-onboarded with complete metadata as follows (replace the example endpoint, model, and sandbox name):

```bash
$$nemoclaw legacy-sandbox destroy
NEMOCLAW_PROVIDER=custom \
NEMOCLAW_ENDPOINT_URL=https://endpoint.example/v1 \
NEMOCLAW_MODEL=your-model-id \
NEMOCLAW_PREFERRED_API=openai-completions \
$$nemoclaw onboard --name legacy-sandbox
```

Hermes deterministically selects `openai-completions` for `compatible-anthropic-endpoint`, so that one route may omit `--inference-api`; explicit incompatible values are rejected.
Use a different `NEMOCLAW_GATEWAY_PORT` when sandboxes need independent routes.
For more information, refer to [Switch Inference Providers](../inference/switch-inference-providers), [NemoClaw CLI Commands Reference](../reference/commands), and [Troubleshooting](../reference/troubleshooting).
- LangChain Deep Agents Code now provides managed `status`, `whoami`, and `identity` commands without launching the interactive UI, validates the installed agent version during onboarding, and keeps credential-shaped or tracing configuration out of persisted runtime metadata.
Its rebuild path validates recreation before destructive handoff and preserves the managed proxy, tool-disclosure, and MCP boundaries.
For more information, refer to [Quickstart with LangChain Deep Agents Code](/user-guide/deepagents/get-started/quickstart), [NemoClaw CLI Commands Reference](../reference/commands), and [Security Best Practices](../security/best-practices).
Expand Down
8 changes: 7 additions & 1 deletion docs/inference/inference-options.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -57,7 +57,7 @@ The managed `dcode` runtime uses Chat Completions through OpenShell even when a
| Google Gemini | Tested | OpenAI-compatible | Uses Google's OpenAI-compatible endpoint |
| Hermes Provider | Hermes only | OpenAI-compatible route | Available when onboarding Hermes Agent through `nemohermes` |
| Local Ollama | Tested with limitations | Local Ollama API | Available when Ollama is installed or running on the host. Validated default models: `qwen3.6:35b` (high VRAM), `nemotron-3-nano:30b` (medium VRAM), `qwen3.5:9b` (low VRAM fallback). |
| Local NVIDIA NIM | Experimental | Local OpenAI-compatible | Requires `NEMOCLAW_EXPERIMENTAL=1` and a NIM-capable NVIDIA GPU. Host must have the NVIDIA Container Toolkit installed and a CDI spec present (`onboard` asserts CDI presence with `assertCdiNvidiaGpuSpecPresent`, `src/lib/onboard/fatal-runtime-preflight.ts`). NIM images pull from `nvcr.io` and require NGC registry login. NemoClaw gates this path behind the experimental flag because it does not auto-select a NIM image for the host today. You must explicitly pick from the validated image list. Managed vLLM has host-specific default models and is not gated on the same boxes. Validated images referenced in `src/lib/inference/config.ts` and `nemoclaw/src/index.ts`: `nvidia/nemotron-3-super-120b-a12b` (default cloud model), `nvidia/nemotron-3-nano-30b-a3b`, `nvidia/llama-3.3-nemotron-super-49b-v1.5`. |
| Local NVIDIA NIM | Experimental | Local OpenAI-compatible | Requires `NEMOCLAW_EXPERIMENTAL=1` and a NIM-capable NVIDIA GPU. Host must have the NVIDIA Container Toolkit installed and a CDI spec present (`onboard` asserts CDI presence with `assertCdiNvidiaGpuSpecPresent`, `src/lib/onboard/fatal-runtime-preflight.ts`). NIM images pull from `nvcr.io` and require NGC registry login. NemoClaw gates this path behind the experimental flag because it does not auto-select a NIM image for the host today. You must explicitly pick from the validated image list. On Linux arm64 DGX Spark and DGX Station hosts, onboarding warns that some NIM images may not publish a `linux/arm64` manifest; the warning is advisory, and the selected image pull can still fail when the registry has no matching platform manifest. Managed vLLM has host-specific default models and is not gated on the same boxes. Validated images referenced in `src/lib/inference/config.ts` and `nemoclaw/src/index.ts`: `nvidia/nemotron-3-super-120b-a12b` (default cloud model), `nvidia/nemotron-3-nano-30b-a3b`, `nvidia/llama-3.3-nemotron-super-49b-v1.5`. |
| Local vLLM (already running) | Tested with limitations | Local OpenAI-compatible | Appears in the onboarding menu when NemoClaw detects a server already on `localhost:8000`. No flag required. Model is whatever the existing server serves. |
| Local vLLM (managed install/start) | Tested with limitations | Local OpenAI-compatible | Appears by default on DGX Spark and DGX Station. Generic Linux NVIDIA GPU hosts require `NEMOCLAW_EXPERIMENTAL=1` or `NEMOCLAW_PROVIDER=install-vllm`. Host must have the NVIDIA Container Toolkit installed and a CDI spec present (`onboard` asserts CDI presence). NemoClaw pulls or starts the stable NGC vLLM container for each host profile. See `src/lib/inference/vllm.ts:55,177` for the pins. DGX Spark and DGX Station use `nvcr.io/nvidia/vllm:26.05.post1-py3`; generic Linux NVIDIA GPU hosts use `nvcr.io/nvidia/vllm:26.03.post1-py3`. Validated defaults are listed in `src/lib/inference/vllm-models.ts`: DGX Spark uses `nvidia/Qwen3.6-35B-A3B-NVFP4`, DGX Station uses `deepseek-ai/DeepSeek-V4-Flash`, and Linux NVIDIA GPU uses `nvidia/NVIDIA-Nemotron-3-Nano-4B-FP8`. Image pulls require NGC registry login (`docker login nvcr.io`); onboard prompts for the NGC API key when authentication is missing. |
{/* provider-status:end */}
Expand Down Expand Up @@ -501,6 +501,12 @@ If the selected vLLM image does not support an argument, the managed container e

NemoClaw can pull, start, and manage a NIM container on hosts with a NIM-capable NVIDIA GPU.

<Warning>
On Linux arm64 DGX Spark and DGX Station hosts, some NIM images may not publish a `linux/arm64` manifest.
NemoClaw prints an advisory warning and still attempts the selected image.
If the registry reports that no matching platform manifest exists, select NVIDIA Endpoints, managed vLLM, or another provider with an image for the host architecture.
</Warning>

Set the experimental flag and run onboard.

```bash
Expand Down
3 changes: 3 additions & 0 deletions docs/inference/use-local-inference.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -266,6 +266,9 @@ Press **Enter** to accept the profile default, or choose a numbered model from t
For scripted runs, set `NEMOCLAW_VLLM_MODEL=<slug>` to choose a registry model without prompting.
If the host reboots and the `nemoclaw-vllm` container is stopped, NemoClaw restarts the managed vLLM container during recovery instead of requiring a fresh onboarding run.
NIM uses the same chat-completions API path restriction as vLLM.
On Linux arm64 DGX Spark and DGX Station hosts, the wizard warns that some NIM images may not publish a `linux/arm64` manifest.
The warning is advisory, so selecting Local NVIDIA NIM still attempts to pull and start the chosen image.
If the registry reports that no matching platform manifest exists, select NVIDIA Endpoints, managed vLLM, or another provider with an image for the host architecture.

On Linux Docker-driver GPU sandboxes, NemoClaw keeps local inference on the OpenShell bridge route and verifies `https://inference.local/v1/models` from inside the sandbox runtime after the sandbox reaches ready.
It treats only a 2xx response as success because that path includes the proxy authentication rewrite the agent uses.
Expand Down
16 changes: 10 additions & 6 deletions docs/reference/commands.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -2029,10 +2029,10 @@ Pass `-f` / `--file <agents.yaml>` to point at the manifest; `--yes` confirms th

</AgentOnly>

### `$$nemoclaw <name> sessions`

<AgentOnly variant="openclaw">

### `$$nemoclaw <name> sessions`

List OpenClaw conversation sessions in the sandbox.
With no subcommand the in-sandbox CLI lists stored sessions for the configured default agent.
NemoClaw invokes `openclaw sessions` via `openshell sandbox exec` and forwards OpenClaw flags verbatim, but filters default list output so internal `nemoclaw-onboard-warmup-*` sessions created during onboarding are hidden from user-facing output.
Expand All @@ -2045,6 +2045,8 @@ $$nemoclaw my-assistant sessions --all-agents --json
</AgentOnly>
<AgentOnly variant="hermes">

### `$$nemoclaw <name> sessions`

List Hermes conversation sessions in the sandbox.
NemoClaw invokes `hermes sessions list` via `openshell sandbox exec`, forwards native Hermes flags such as `--source` and `--limit`, and streams the output unchanged.

Expand All @@ -2055,10 +2057,10 @@ $$nemoclaw my-assistant sessions --source cli --limit 20

</AgentOnly>

### `$$nemoclaw <name> sessions list`

<AgentOnly variant="openclaw">

### `$$nemoclaw <name> sessions list`

Invoke `openclaw sessions list` inside the sandbox.
NemoClaw forwards every flag the in-sandbox CLI accepts (`--agent`, `--all-agents`, `--active`, `--limit`, `--json`, `--store`, `--verbose`) and filters the resulting default table or JSON so internal `nemoclaw-onboard-warmup-*` sessions are hidden.

Expand All @@ -2070,6 +2072,8 @@ $$nemoclaw my-assistant sessions list --agent work --json
</AgentOnly>
<AgentOnly variant="hermes">

### `$$nemoclaw <name> sessions list`

Invoke `hermes sessions list` inside the sandbox.
NemoClaw forwards native Hermes flags such as `--source` and `--limit` and streams the output unchanged.

Expand Down Expand Up @@ -2810,9 +2814,9 @@ $$nemoclaw credentials reset nvidia-prod
### `$$nemoclaw gc`

Remove orphaned sandbox Docker images from the host.
Each `$$nemoclaw onboard` builds an `openshell/sandbox-from:<timestamp>` image (~765 MB).
Sandbox creation can build images in the gateway-managed `openshell/sandbox-from` repository or the locally prebuilt `nemoclaw-sandbox-local` repository.
The `destroy` and `rebuild` commands clean up the image automatically, but images from older NemoClaw versions or interrupted operations may remain.
This command lists all `openshell/sandbox-from:*` images, cross-references the sandbox registry, and removes any that are no longer associated with a registered sandbox.
This command lists images from both repositories, cross-references the sandbox registry, and removes any that are no longer associated with a registered sandbox.

```bash
$$nemoclaw gc [--dry-run] [--yes|-y|--force]
Expand Down
Loading
Loading