EN | DE
A model-agnostic computer-use core: one agent loop, any reasoning model behind a single interface.
open-compute is a small, dependency-light Python core for building computer-use
agents (LLM-driven GUI / desktop / browser automation). It implements the
perception → model-tool-call → action → feedback loop and keeps the
reasoning model swappable behind a single ComputerBackend interface. No
provider is privileged: Anthropic Claude and OpenAI CUA are two equally-ranked
API backends, and the offline mock backend is the default. A keyless path
also exists today via Mode A, where the host model itself reasons — and it can
run that loop either inline or in a self-spawned subagent for context economy
(see usage pattern). The core has
zero runtime dependencies; vendor SDKs (anthropic, openai) are
optional, lazily imported extras — import open_compute works with none of
them installed, and the default mock wiring runs fully offline.
Note
AI / LLM Integration Notice: open-compute includes a machine-readable llms.txt file designed for AI agents, RAG crawlers, and LLM-assisted workflows.
Every computer-use model — Anthropic's Claude computer tool and OpenAI's
computer-use tool — shares the same agent-loop shape but differs in transport,
coordinate frame, and action names. open-compute factors out the common parts so
you write the loop once and swap the reasoning model freely behind one
ComputerBackend interface:
- A canonical action schema with one mapper per backend.
- Normalized (0..1) coordinates internally, denormalized per backend / resolution / DPI in one tested utility — the DPI problem solved centrally.
- A central safety gate ("confirm before risky actions") evaluated before every action.
- A hybrid perception interface (screenshot + Set-of-Marks / accessibility / DOM), so you can move from pure pixel-vision to semantic targeting later.
+-----------------------------------------+
| AGENT LOOP / ORCHESTRATOR |
| goal -> perceive -> backend -> safety |
| -> execute -> re-perceive |
+-------------------+---------------------+
|
+-----------------------------------+-----------------------------------+
| | |
+-------v---------+ +----------v-----------+ +----------v----------+
| PERCEPTION | | CANONICAL ACTIONS | | SAFETY / POLICY |
| - screenshot | | click/type/key/ | | - confirm-at-action |
| - set-of-marks | | scroll/drag/wait/ | | - allow / deny list |
| (OmniParser)* | | screenshot + OS ext | | - read-only mode |
| - accessibility*| | (launch/activate) | | - audit log |
+-------+---------+ +----------+-----------+ +----------+----------+
| | |
+-----------------+-----------------+----------------------------------+
|
+-----------v------------+ COORDINATE / DPI NORMALIZATION
| BACKEND ABSTRACTION | - internal: normalized (0..1)
| (ComputerBackend) | - denormalize per backend:
+-----+--------+---------+ * Claude: global px (display_w x display_h)
| | | * OpenAI: px (computer_call)
+-----------+ | +-----------+ * Mock: synthetic
| | |
+-------v-------+ +--------v-------+ +-----v---------+
| Claude | | OpenAI CUA | | Mock backend |
| computer_2025 | | computer-use- | | (no SDK, |
| 1124 + beta | | preview [?] | | offline) |
| (host runs) | | (host runs) | | |
+---------------+ +----------------+ +---------------+
* = stub / interface in this release (see Status)
Important
Not on PyPI — install from Git. This project has no PyPI release yet. The
name open-compute on PyPI is taken by an unrelated project ("multi-agent
systems for healthtech"), so a plain pip install open-compute installs
someone else's package. Always install from this repository:
pip install "git+https://github.com/ellmos-ai/open-compute.git" # core only, zero runtime deps
pip install "open-compute[claude] @ git+https://github.com/ellmos-ai/open-compute.git" # + anthropic SDKThe same extra @ git+… form works for every extra below:
| Extra | Adds |
|---|---|
claude |
anthropic SDK |
openai |
openai SDK |
local |
mss — real Windows screenshots + input |
wgc |
WGC fallback for DirectX surfaces (pulls numpy/OpenCV) |
compose |
Pillow — Before|After composite + annotated shots |
watch |
watchdog — native FS events for the directory-watch feed |
clirec |
external clirec package for oc rec workflows |
record |
clirec[record] capture backend compatibility |
mcp |
mcp SDK — MCP server (console script: open-compute-mcp) |
dev |
pytest |
all |
anthropic, openai, playwright, mss, WGC, Pillow, watchdog, clirec, mcp |
Extras combine as usual, e.g. open-compute[local,wgc,claude]. Working from a
clone instead? pip install -e ".[local,claude]" from the repository root.
Until clirec has a package release, install it directly when using oc rec:
pip install git+https://github.com/ellmos-ai/clirec.gitPython 3.10+.
Run oc capture / oc do manually from a Claude Code session. The session
model sees the PNG via the Read tool and decides the next action:
# 1. Install the local extra (Windows only; provides real screenshots + input)
pip install "open-compute[local] @ git+https://github.com/ellmos-ai/open-compute.git"
# 2. Capture a screenshot — saved automatically to _session/ (never loose on Desktop)
oc capture
# -> {"path": ".../_session/0001_20260620_143200.png", "width": 1920, "height": 1080}
# Then: read the PNG with your Read tool to see the screen.
# 3a. Execute one canonical action (single, backwards-compatible)
oc do '{"type":"mouse_move","x":0.5,"y":0.5}' --mode allow_all
oc do '{"type":"left_click","x":0.25,"y":0.1}' --yes # --yes = agent pre-approved
# 3b. Execute with Before|After composite (Pillow optional)
oc do '{"type":"left_click","x":0.5,"y":0.3}' --label "click_ok" --yes
# -> {"result":"executed","action":"left_click","composite":"_session/0002_click_ok.png"}
# 3c. Execute a batch/macro (JSON array, one call = multiple actions)
oc do '[{"type":"mouse_move","x":0.5,"y":0.5},{"type":"left_click","x":0.5,"y":0.3}]' --yes
# -> {"result":"batch","count":2,"width":1920,"height":1080}
# 3d. Ensure the target window is in the foreground before acting
oc do '{"type":"left_click","x":0.5,"y":0.3}' --ensure-foreground "Word" --yes
# 3e. Save a full-res after-shot + annotated click marker (v0.5, Pillow optional)
oc do '{"type":"left_click","x":0.5,"y":0.3}' --yes --fullres
# -> {"result":"executed",...,"fullres_annotated":"_session/...fullres.png"}
# 3f. Capture only the active window's bounding rect (v0.5, Windows)
oc capture --window "Word"
# -> {"path":"...","width":800,"height":600,"window":"Word","region":{...}}
# 3g. Watch a directory for changes (v0.5)
oc watch-dir ~/Downloads --for 5 # collect 5 s, print JSON events
oc watch-dir ~/Downloads --once # one-time snapshot diff
# 3h. Explicit companion handoff (mutations need a granted, scoped lease)
oc session companion --owner local-user
oc session request-control --owner agent-a --scope window:42 --ttl 60
oc session grant --lease-id <lease_id-from-previous-output>
oc window minimize --hwnd 42 --yes
# 3i. Bounded, deduplicated window capture (full screen needs explicit opt-in)
oc capture-series --window "Word" --max-frames 8 --stable-frames 2
# 4. Recapture and repeat until done (or read the "composite" After-shot directly).See SKILL.md for the full loop protocol, action schema, coordinate guide, and
environment variable reference.
The backend is selected by name; claude and openai are equally supported
(each needs its own key + extra). For a keyless path, use Mode A above — the
host model reasons itself, optionally in a self-spawned subagent (see
usage pattern).
# Claude (needs ANTHROPIC_API_KEY + open-compute[local,claude]):
oc run "Find the latest invoice in the Downloads folder" --backend claude --max-steps 15
# OpenAI (needs OPENAI_API_KEY + open-compute[local,openai]):
oc run "Find the latest invoice in the Downloads folder" --backend openai --max-steps 15Or in Python — get_backend(name, ...) builds whichever you name; inject your
own executor or use LocalExecutor:
from open_compute import AgentLoop, Config, get_backend
from open_compute.drivers.local import LocalExecutor # Windows; needs mss
from open_compute.safety import SafetyPolicy
executor = LocalExecutor() # real display + input
config = Config(backend="claude", scope="os",
display_width=executor.width, display_height=executor.height)
backend = get_backend("claude", executor.width, executor.height, model="claude-opus-4-8")
loop = AgentLoop(
config,
backend=backend,
executor=executor,
policy=SafetyPolicy(mode="confirm",
confirm_callback=lambda a: input(f"run {a.type.value}? [y/N] ") == "y"),
)
loop.run("Find the latest invoice in the Downloads folder")from open_compute import AgentLoop, Config
loop = AgentLoop(Config(backend="mock", safety_mode="allow_all"))
result = loop.run("Open the settings page and enable dark mode")
print(result.done, result.steps)
for trace in result.traces:
print(trace.index, trace.backend_message, [a.type.value for a in trace.executed])Expose the keyless Mode A loop to any MCP client as native tools — the
client is the reasoner (no API key, model-agnostic). Versus driving oc by hand,
a long-lived server keeps one warm LocalExecutor resident (no Python restart
per action) and returns screenshots as MCP image blocks. Windows-only for real
capture/input.
pip install "open-compute[mcp,local,uia,wgc] @ git+https://github.com/ellmos-ai/open-compute.git"
open-compute-mcp # stdio server (console script)Tools: capture · do (single or batch canonical actions) · tree ·
click_name · invoke (UIA semantic targeting) · list_windows ·
get_screen_size · watch_dir · push_status · rec_replay · signal_show /
signal_hide / signal_status / signal_abort (human-in-the-loop screen
signal) · chat · talk (push-to-talk). Coordinates are
normalized 0..1; list_windows and get_screen_size describe that frame, so the
client can name a window exactly instead of guessing a title substring.
Hardware-composited windows (wgc extra). A GDI grab of a DirectX window —
Roblox Studio, Blender, a GPU-accelerated browser — does not fail; it quietly
returns an all-black rectangle. capture(window=...) therefore checks the
frame and, when it comes back blank, re-grabs it through Windows.Graphics.Capture.
Install open-compute[wgc] for that fallback; without it a black frame is still
returned rather than failing the call. OC_WGC_WINDOWS (comma-separated title
substrings) skips the GDI attempt outright for windows known to need WGC.
Note that WGC only produces a frame when the window redraws: an idle or
non-capturable window fails fast (bounded, a few seconds) instead of hanging.
Capture budget (token cost). A vision model is billed per pixel, so a full-HD
capture is by far the most expensive thing this server returns — and every frame
stays in the conversation, so the cost is paid again on each following request.
Because all coordinates here are normalized 0..1, shrinking the image costs
nothing in control accuracy; only legibility drops. Three knobs:
| Variable | Effect | Cost of a 1920×1080 grab |
|---|---|---|
| (unset) | full resolution | ~1600 tokens |
OC_CAPTURE_SCALE=0.5 |
halve both edges | ~690 tokens |
OC_CAPTURE_MAX_DIM=768 |
cap the longest edge | ~440 tokens |
OC_CAPTURE_GRAYSCALE=1 |
drop colour | payload only — not tokens, which follow pixel count alone |
OC_CAPTURE_SCALE=0.5 is the sweet spot for GUI work: buttons and field borders
stay clearly identifiable, only small body text gets hard to read. Both size knobs
compose (scale first, then the cap), and a failure to shrink never fails the
capture — the original frame is returned instead.
Safety. OC_SAFETY_MODE is an operator ceiling (confirm default ·
read_only · allow_all); a per-call mode can only tighten it, never loosen it,
so a prompt-injected agent cannot escape a read_only/confirm server via
mode="allow_all". Because stdio MCP has no server→client confirm callback,
confirm/read_only return a needs_confirmation/deny result without acting.
For interactive use, run the server with OC_SAFETY_MODE=allow_all in an isolated
VM and let the client's tool-permission dialog be the human-in-the-loop. Optional
OC_DENY (comma-separated action types) is a hard deny list.
Auto-signal (OC_SIGNAL_AUTO). Set it to a SessionMode name (e.g.
control) to auto-show the screen-usage overlay the first time a
state-changing tool (do / click_name / invoke / rec_replay) actually
passes the safety gate — no separate signal_show call to remember before
the model starts steering. It never overrides an already-visible signal
(manual or auto, any mode) and never fires from a gate-blocked call or a
read-only tool. Unset or off (the default) disables it; an invalid mode
name surfaces as auto_signal_error in the tool result instead of failing
the call. See signal_show/signal_hide/signal_status below for the
manual controls and OC_SIGNAL_CONFIG for per-mode colors.
Auto-hide (OC_SIGNAL_IDLE_HIDE). An auto-shown overlay takes itself down
once the steering stops: every state-changing tool call re-arms an idle
countdown, and when it expires with no further action the overlay is hidden.
The value is seconds, default 60; 0, an empty value, or off disables the
auto-hide and keeps the overlay up until signal_hide (the pre-0.7 behavior).
Only an overlay that OC_SIGNAL_AUTO put up is ever swept away — one you asked
for with signal_show stays until you hide it, and a manual signal_show over
an auto-shown overlay takes ownership and cancels the countdown. signal_status
reports both (auto_shown, idle_hide_armed); an unusable value surfaces as
signal_idle_hide_error in the tool result instead of failing the action.
Troubleshooting: do/click_name only ever return needs_confirmation and never
act. That is the confirm ceiling working as designed under stdio MCP — there is
no confirm callback, so the server reports instead of acting. Fix for interactive
use: set "env": {"OC_SAFETY_MODE": "allow_all"} in the server registration and let
the client's tool-approval dialog gate each action (do not auto-allow the
do/click_name/invoke tools there, or you lose that gate). Note that the env
change only takes effect when the server process (re)starts — an already-connected
client keeps the old ceiling until it reconnects.
Client config (via uvx, no manual install):
{ "mcpServers": { "open-compute": {
"command": "uvx",
"args": ["--from", "open-compute[mcp,local,uia] @ git+https://github.com/ellmos-ai/open-compute.git", "open-compute-mcp"] } } }The snippet above starts in the safe confirm ceiling — the server reports
actions but does not perform them. To let it act, add
"env": {"OC_SAFETY_MODE": "allow_all"} (isolated VM), gated by the client dialog.
An npm launcher (npx open-compute-mcp) is also published for parity with Node MCP
servers and is listed in the Glama MCP directory.
The MCP server is the ideal shape for short, inline tasks; for long,
context-heavy runs, still delegate to a self-spawned subagent (see the usage
pattern below) and call these tools inside it.
| Backend | SDK | Tool / model | Coordinates | Status |
|---|---|---|---|---|
mock |
none | scripted, offline | synthetic | Fully implemented (default backend) |
claude |
anthropic (lazy) |
computer tool computer_20251124, beta header computer-use-2025-11-24, default model claude-opus-4-8 |
global pixels; host executes | Implemented; tested via injected client |
openai |
openai (lazy) |
computer-use, model computer-use-preview (configurable, [UNSICHER]) |
pixels; host executes | Implemented; model name / request shape not fully verified |
local (foreign reasoner) |
none | a different model as reasoner — local Ollama, or agy / codex / kimi CLIs | host executes | Separate, low-priority, optional idea — would be a real new backend with possible capability differences. Not scheduled. |
The keyless / no-API path is not a backend row — it is Mode A, where the host model itself reasons (inline, or in a self-spawned subagent for context economy; see usage pattern).
The implemented backends (mock / claude / openai) share one
ComputerBackend Protocol and are dispatched by name from get_backend()
(open_compute/backends/factory.py) — no provider is hard-wired into the loop.
The Claude tool type / beta header pair
is configurable on the backend
(tool_type=, beta_header=) so you can target the older computer_20250124
/ computer-use-2025-01-24 pair on older models.
| Executor | Requires | Platform | Status |
|---|---|---|---|
MockExecutor |
none | any | Fully implemented; used in tests and dry-runs |
LocalExecutor |
mss (open-compute[local]), optional WGC fallback (open-compute[wgc]) |
Windows only | Implemented; oc capture live-tested (368 KB PNG at 1920×1080); oc do mouse_move live-tested |
Fully implemented and tested
-
Canonical action schema +
to_claude/to_openaimappers. -
Coordinate normalize / denormalize / rescale.
-
Safety policy gate (
confirm/allow_all/read_only, deny list, confirmation callback, audit log). -
Configdataclass + JSON loader. -
Agent loop orchestrator (dry-run via mocks).
-
Headless cooperative core (
cooperative.py,human_activity.py): injectable perceive/stabilize/act/verify ports, scoped lease and human/emergency-stop gates, no-replay action IDs, bounded retries, screen-prompt-injection blocking, hash-chained sanitized audit, explicit retention/deletion and crash cleanup.GetLastInputInfois a single-shot adapter tested only with injected callables; no hook or monitor is enabled. -
Backend dispatch via factory +
MockBackend; Claude backend tested with an injected fake client. -
LocalExecutor(Windows,open-compute[local]): real screenshot via mss, real mouse/keyboard via ctypes SendInput with VIRTUALDESK + DPI-awareness. Optionalopen-compute[wgc]adds a Windows.Graphics.Capture fallback for DirectX / hardware-composited surfaces when mss/GDI capture fails. Action dispatch for all action types. Live-tested:oc capture→ PNG 368 KB (1920×1080);oc do mouse_move→ cursor moved. -
ocCLI (oc capture/oc do/oc run): Mode A (no-key skill loop) and Mode B (autonomous AgentLoop with API backend) wired end-to-end.- v0.3:
oc capturedefaults to_session/(never loose in CWD/Desktop). - v0.3:
oc doaccepts JSON arrays (batch/macro) and--labelfor automatic Before|After composite screenshots. - v0.3:
--ensure-foreground SUBSTR/OC_ALWAYS_FOREGROUNDonoc doandoc runfor automatic window activation before actions. - v0.3:
Config.always_foregroundfield +[compose]optional extra (Pillow).
- v0.3:
-
SKILL.md: loop protocol for the session-agent (Mode A). -
Multi-feed abstraction (v0.4,
open_compute/feeds/):PerceptionFeed+Targeterprotocols,ScreenshotFeed(pixel), and a runtime feed registry (available_feeds()) with graceful capability detection. -
UiaWindowsFeed(v0.4, Windows,open-compute[uia]): UIA element-tree perception + semantic targeting.observe()walks the ControlView tree;resolve()does exact > prefix > contains disambiguation;invoke()does click-free activation via InvokePattern → Toggle → SelectionItem → LegacyIAccessible fallback.center_normis the exact inverse ofLocalExecutor's virtual-desktop mapping (round-trip covered by tests, incl. negative multi-monitor origin). The full invoke/resolve/coordinate logic is unit-tested withuiautomationmocked; real-OS smoke tests (oc tree,oc click-name --mode confirm,oc invoke --mode confirm) were run on Windows 11 — seeCHANGELOG.md. -
ocCLI (v0.4):oc tree,oc click-name,oc invoke— all routed through the Safety gate. -
DirwatchFeed(v0.5,open_compute/feeds/dirwatch.py): directory-watch event feed. Monitors configured paths and emits change events (created / modified / deleted / moved) into a rolling deque. Two backends: watchdog (MIT, native OS events —open-compute[watch]) or stdlib polling (always available without extras).available()always returnsTrue.oc watch-dir <path> [--for SECS] [--once]CLI. -
Full-res / annotated verification shot (v0.5):
oc do --fullresandoc click-name --fullressave an additional full-resolution after-shot alongside the composite. Pillow (optional) annotates the click position with a red circle + crosshair. JSON keys:"fullres"/"fullres_annotated". -
oc capture --window SUBSTR(v0.5, Windows): captures only the bounding rect of the named window via Win32GetWindowRect. Case-insensitive, whitespace-normalized substring match (same convention asUiaWindowsFeed). -
FeedManager(v0.6,open_compute/feed_manager.py): dosierte Push-Auto-Injektion. Collects available feeds, applies change-detection per cycle (State-Feeds: SHA-256 hash; Event-Feeds: rolling window), dispatches to anInjectorSink. Dosage modes per feed:full|delta|notify|off; runtime-adjustable viaset_dosage().LocalFileInjector(working default; writes to_state/inject_queue/).BachInjectorAdapter(stub; seefeed_manager.pydocstring for activation instructions).oc push --status/oc push --onceCLI. -
LearningManager(v0.6,open_compute/learning.py): Bandit/Bayes weighting (BetaPrior), use-case profiles (JSON, warmstart viaapply_profile_to_manager()), and cross-session LESSONS-LEARNED (JSONL). All state in gitignored_state/.
Interface / stub (honest)
- Browser driver and OS driver are interfaces only (no Playwright / CDP / host implementation yet).
- Perception providers other than
ScreenshotPerceptionand the v0.4 UIA feed (Set-of-Marks, OCR, vision overlays, DOM) are not yet implemented. BachInjectorAdapteris a documented stub;LocalFileInjectoris the working default sink.- Always-on push daemon (permanent background loop) is not yet implemented.
- Live human-input monitoring, ownership-overlay rendering, global emergency hotkeys, voice, virtual-display/session control, and any productive wiring of the headless cooperative core are not implemented or activated.
oc recis a lazy compatibility shim for the externalellmos-ai/clirecpackage; installclireconly when recording/replay workflows are needed.- The UIA feed is Windows-only; Linux (AT-SPI) and macOS (AXUIElement) accessibility feeds are open / planned.
- The OpenAI backend's model name and exact Responses-API request shape are not fully verified — validate against live OpenAI docs before production.
- The self-subagent mode (b) is a usage pattern (docs), not new reasoning code — see below. A foreign / local reasoner (Ollama / agy / codex / kimi) is a separate, low-priority, optional idea, not implemented.
See TODO.md for the full breakdown.
Pattern, not a new backend. Same host model, no API key — only the context budget differs. Full design in
ARCHITECTURE.md("Host-Modell-Kontext: Inline (a) vs. Selbst-Subagent (b)").
When the host model (e.g. Claude Code on a subscription) runs the no-key Mode A loop, it can spend its context two ways — same model, same vision, same reasoning:
- (a) Inline (today's solution). The host model runs
capture → decide → do → recapturein its own context. Best for short / simple tasks (a few steps). - (b) Self-subagent (concept). The host model spawns a subagent of
itself (e.g. via a
Task) that runs the whole loop in the subagent's context and returns only the distilled result ("invoice found at …"). The main context stays clean; it "feels like API" but is the same model — the win is context economy, not a reasoning/vision trade-off. Best for long / repeated / context-heavy tasks.
The model decides per task, exactly like normal subagent delegation. Rough heuristic: short → inline (a); long / repeated / context-heavy → spawn a subagent (b).
A persistent 24h experience-subagent is an optional variant of (b): a
long-lived self-subagent that takes repeated jobs and reuses accumulated
experience via the existing learning.py (BetaPrior / use-case profiles /
LESSONS-LEARNED in _state/). Experience lives in _state/ (persistent), not
in the volatile subagent context. (Lessons should carry decay / confidence to
avoid false lessons — a small additive change, not yet implemented.)
A different model as reasoner (local Ollama, or agy / codex / kimi CLIs) is a separate, low-priority, optional idea — that would be a real new
ComputerBackendwith possible capability differences, and is not mode (b).
Computer-use is powerful. The default SafetyPolicy mode is confirm: clicks,
typing, key presses, drags, and app launches are blocked unless a confirmation
callback approves them. Recommended practice (mirrors both vendors' guidance):
- Run real backends in an isolated VM or container, never your main desktop.
- Keep a human in the loop.
- Treat on-screen content as untrusted (prompt-injection risk).
See SECURITY.md.
Generated discovery projection for module:open-compute from catalog:v4-bundles (546290dafbaafd810df1d59ef5a3d7183738472b48cd5a8a81f1e8f2b64d852e).
Target repository visibility: public. Bundle manifests remain the membership authority; this section does not install or activate components.
Discovery approval: public module-registry record, explicit default-deny bundle allowlist.
- Bundle recipe visibility:
private; role:declared-component; requirement:recommended. - module partners:
module:ai-media-editor,module:report-forge,module:web-scraper. - skill partners:
skill:textproduction,skill:video-transcriber.
- Bundle recipe visibility:
private; role:declared-component; requirement:recommended. - module partners:
module:ApiProber,module:clirec,module:connectors,module:software-endpoint-registry. - skill partners:
skill:ai-portable-setup.
Composition and runtime details are intentionally omitted.
python -X utf8 -m pytest -qTests are mock-only and require no SDK; pip install -e ".[dev]" from a clone
installs pytest. Current full-suite state: 551 passed, 1 skipped (2026-08-06).
MIT — see LICENSE.
