Skip to content

Hands-Free Fleet Orchestration

Computer Use and Voice

Speak a goal and hand it to an agent on your own machine. The design keeps every action visible and puts the approval switch on your phone. Downloads are not live yet, and detailed controls will be documented at launch.

How the Workflow Operates

STAGE 01

Speak the Goal

Local transcription, one command

You say what you want, such as 'Run the staging checks on the checkout flow.' The voice tool is designed to transcribe your speech locally with whisper.cpp and send the text as a command to one of your harnesses.

STAGE 02

The Agent Does the Work

Your harness, your tools

The harness hands the command to an agent on your own machine, with the tools and access you chose for it. How computer-use agents read and act on the screen will be documented at launch.

STAGE 03

Show Every Action

Visible steps, not hidden work

The design goal is simple: every action an agent takes is shown to you as it happens. The exact interface will be documented at launch.

STAGE 04

Independent Verification

A different vendor reviews the work

When the task finishes, a verifier from a different vendor reviews the result. A second perspective can catch what the first agent missed. It is not a guarantee that every task is right.

Safety Model

Human Oversight by Design

Autonomous work needs clear boundaries. These are the design goals for computer use and voice. Detailed controls will be documented at launch.

Every Action Shown

Designed so agents never work out of sight. Each action is meant to be shown to you as it happens.

Approve from Your Phone

Planned: approve once, for this session, or always, from your phone. Which actions ask for approval will be documented at launch.

Independent Review

Every agent task gets a verifier from a different vendor. It is a second check, not a guarantee.

Documented at Launch

Controls for pausing or stopping an agent, and how its work is kept apart from yours, will be documented at launch. We do not describe controls before they ship.

Honest Engineering Boundaries

Capabilities and Real Technical Limits

Computer use and voice are useful for structured work, and they have real limits. Here they are, plainly.

CLI and APIs First

Command-line tools and direct APIs are usually faster and more predictable than screen-based automation. When a job can be done through a CLI or an API, that is usually the better route.

Latency and Visual Reasoning

Vision models take seconds to read each screen. Computer use suits background procedures and end-to-end checks, not real-time control.

OS Permission Requirements

Screen capture and input control need operating system permissions that you grant. There is no zero-config way around the operating system security model.

Credentials and Enterprise Tokens

Planned managed fleets will run only on your own API keys or enterprise tokens. We never run agents on customers' personal subscription seats from a hosted backend.

Frequently Asked Questions

How will phone approval work?

Phone approval is planned, not shipped. The design: when an agent reaches an action that needs your say, you approve once, for this session, or always. Which actions ask, and how your phone is linked, will be documented at launch.

Does my voice audio leave my local machine?

The voice tool is designed to transcribe speech locally with whisper.cpp, so the audio stays on your machine. Only the text command goes to your harness. What happens next depends on the agent and model you chose for that harness.

Can an agent take over my computer while I am working?

Computer use is not released yet. How agents share a machine with you, and how to keep their work apart from yours, will be documented at launch.