Tool belt
Tools agents can actually use
MCP servers, plugins, APIs, native app control, and internal workflow integrations.
Selected work
Autonomous departments, agent products, on-device systems — live and in real use.
What we mean by harness
Founders do not need a smarter demo. They need an AI system that can touch real workflows without corrupting state, skipping approvals, leaking data, or running up cost in silence.
Agent operating system
Tool belt
Control
State
Judgment
Visibility
Recovery
A prompt can answer. A harness can safely read, write, remember, recover, and explain its work.
Tool belt
MCP servers, plugins, APIs, native app control, and internal workflow integrations.
Control
Read-versus-write boundaries, scoped credentials, approval gates, and audit trails.
State
Checkpoints, shared files, durable context, and exact procedural memory where embeddings are wrong.
Judgment
The right model for the task, with validators, routing rules, and fine-tuning only where it earns its keep.
Visibility
Traces, per-output cost, retry budgets, SLIs, alerts, and the ability to explain what happened.
Recovery
Idempotency, circuit breakers, resumable workflows, and failure modes designed before launch.
Featured shipped system
The point is not that there are agents. The point is that they have tools, state, gates, traces, and recovery paths.
One brief becomes finished, scheduled, multi-platform content through a governed agent tree.
Fourteen named specialists are only the first layer. The workflow fans into depth-4 leaf agents for research, drafting, QA, packaging, routing, and publishing. The durable piece is the harness: shared file bus, checkpointed state, idempotent dispatch, eval-gated routing, cost tracing, and human approval gates.
What it proves
depth-4 agent tree
What this prevents
CEOs do not buy checkpoints. They buy confidence that an AI workflow will not rerun a payment, skip approval, lose state, or become impossible to debug when volume rises.
Prompt + tool call
Tool runs
email, payment, row, ticket
Worker dies
after the side effect lands
Workflow retries
without knowing what happened
Business state corrupts
quietly, at production speed
This is how a working demo becomes an operations problem.
Harnessed agent
Checkpoint written
state before side effects
Bounded tool runs
permissioned + idempotent
skip duplicateFailure is traced
retry budget and cost guard
traceSystem resumes
no duplicate action
resumeThe harness makes failure recoverable by construction.
tool_call = agent_id + step + payload_hash
Case studies
A natural-language agent that drives a native real-time engine, in paying users' hands.
A full-stack AI product built on a single-coordinator, multi-tool agent: a conversational interface wired to a native real-time engine, with subscriptions, cloud infrastructure, and end-to-end observability. Shipped across web and desktop, in production.
What it proves
Clinical-grade transcription where the audio never leaves the machine.
A HIPAA-adjacent capture system, compliance from day one. Transcription runs on-device with local ML, so sensitive audio never leaves the machine — a custom low-latency audio driver, macOS System Extensions and XPC, a sandboxed, crash-isolated build. Shipped.
What it proves
Messy specs in, schema-validated data out — every field citing its source page.
Spec, quote, and proposal PDFs and DOCX in; schema-validated fields out, each citing the source page. Layout-aware parsing preserves tables, and output is enforced through the tool-use API — no fragile JSON parsing — then handed to an editable review UI with Excel and PDF export.
What it proves
Bring us in when
These are buying moments, not service categories. They usually happen after the demo works and before the system is safe enough to become part of the business.
Buying moment
A whole function should run from one business input, with agents decomposing the work and humans approving the output.
Buying moment
It works until an API times out, a tool runs twice, costs spike, or nobody can explain why the workflow made a decision.
Buying moment
The model needs a real product around it: auth, billing, data, APIs, native integrations, background jobs, admin surfaces, and deployment.
Buying moment
PDFs, audio, specs, quotes, and operational documents need to become schema-valid outputs your team can inspect, cite, approve, and export.
Open-source infrastructure
MCP servers and agent tooling we build in the open. Public, installable, used outside our work.
An MCP server giving AI agents live control of a target application. FastMCP, on PyPI.
An open-source Claude Code plugin for contract-driven, multi-agent goal execution. A subagent judge gates completion against an explicit Definition of Done, so a passing validator can never ship a placeholder.
A 40-tool desktop-automation MCP server (TypeScript, macOS). We contribute rather than author it — landing macOS stability, GPU detection, and accessibility fixes upstream.
An OpenClaw plugin, published on ClawHub, that lets agents search live events through the Ticketmaster Discovery API.
An OpenClaw plugin, published on ClawHub, that teaches agents to drive AI video production — explainers, trailers, motion graphics — through a local OpenMontage pipeline.
Where this work has gone
Real proposals and shipped engagements.
We will tell you honestly whether we are the right team.