Blog · Computer use · 08 Oct 2026
Blog
Computer use

Computer use agent:What it operates,what it may do

An agent that clicks can reach software that never had an API. Which tasks fit, which permissions it gets, and where a person has to say yes.

Get in touch
Enquiry
Nikolai Schöbel und Jeremias Burger, Co-Founder Scalableloops

Let's talk about your project.

First we check whether the project fits your business model. Then you get a proposal with phases and effort.

Have a first task reviewed or call: +49 151 1576 5566
Blog · Computer use

A computer use agent operates software the way a person does: it looks at a screenshot, decides the next click or keystroke, carries it out, then looks at the next screenshot. Because it works through mouse and keyboard, the software needs no API. That is what makes it interesting for small and mid-sized companies, because an ageing ERP system, a supplier portal or a government form can finally be automated without asking the vendor for an interface. OpenAI added the capability to its Agents API on 29 September 2026; at Microsoft and Anthropic it is already generally available. And all three vendors write the same caveats in their own documentation: the agent belongs on a dedicated, isolated machine, it gets only the permissions the task requires, and a human has to approve anything that cannot be undone.

In brief
  • Computer use means mouse and keyboard instead of an API: look at a screenshot, click or type, check the result, repeat. The agent sees the same interface your staff see.
  • The value sits with software that has no API. Microsoft puts it plainly in its own documentation: if a person can use an app or website, computer use can too.
  • Permissions decide the risk, not the model. OpenAI, Microsoft and Anthropic independently require a dedicated machine, an allow list of permitted sites and apps, and a low-privilege account.
  • The built-in request for human input is not a safeguard. Microsoft states explicitly that you should not rely on the system always asking a person before it proceeds.
Published 8 Oct 2026Nikolai Schöbel and Jeremias Burger9 min read
Nikolai SchöbelJeremias Burger

Nikolai Schöbel and Jeremias Burger

Co-founders of Scalableloops GmbH. Nikolai Schöbel leads online marketing and AI strategy, Jeremias Burger the AI architecture. Both build AI systems and train teams on them in their own agency work.

On this page
  1. What does "computer use agent" mean?
  2. What OpenAI, Claude and Copilot Studio change for computer use agents
  3. Which screen tasks suit a computer use agent
  4. What permissions a computer use agent actually needs
  5. Where a person must approve before a computer use agent submits anything
  6. When a computer use agent pays off for your business
  7. Where computer use agents still fall short, from prompt injection to Citrix
  8. Frequently asked questions
  9. How to test a computer use agent on a first task
  10. Where the information on this page comes from
Basics

What does "computer use agent" mean?

A computer use agent runs a loop that is easy to follow. It receives a task in plain language, say transferring invoice data from a PDF into an entry form. It then sees a screenshot of the machine, picks the next action, performs it, and asks for the result as a fresh screenshot. OpenAI describes the pattern in its own guide as send task, execute returned actions, return screenshot.

The actions themselves are exactly what a person does with mouse and keyboard. OpenAI's documentation lists click, double click, drag, move, scroll, keypress, type, wait and screenshot. Anthropic ships a longer set of seventeen member tools, including right click, triple click, zooming into a region at full resolution and holding a key for a set duration. That is the whole vocabulary, and that is the appeal: software a person can operate by hand can be operated this way too.

The contrast with ordinary automation is therefore sharp. An API exchanges data in a fixed format. It is fast and dependable, but the vendor has to offer it. A computer use agent goes through the interface instead, so it needs nothing from the vendor, but it is slower and less dependable. Where an API exists, the API stays the better choice. What an agent is in general, and what separates it from a chatbot, is in our piece on the AI agent.

What changed

What OpenAI, Claude and Copilot Studio change for computer use agents

Computer use agents became part of OpenAI's Agents API on 29 September 2026, announced at the company's developer conference. The OpenAI developer forum carries the line that computer use in the Agents API lets agents operate software through its UI. InfoQ summarised it on 2 October 2026 by saying applications can now operate software through graphical interfaces, and that the functionality is available through the API and in Codex and ChatGPT Work for selected plans.

A new model came with it. According to both sources, GPT-6.1 Sol improves coding and computer use, and its standard token prices sit at one fifth of those for GPT-6 Astra. OpenAI's own guide to the computer tool names that model as the one to use. On availability, though, the two sources disagree: the forum summary says the model is coming soon, while InfoQ describes it as usable through the API. We therefore give no launch date and report only what the guide states.

In practice the shift matters more than the model. Until now you had to build the screenshot, action and verification loop yourself, or rent a finished platform. Now it sits inside the API teams already use. Falling prices help as well, because a screen task costs many small steps, and every one of them carries a screenshot.

29 Sep

is when OpenAI brought computer use into the Agents API. GPT-6.1 Sol costs one fifth of GPT-6 Astra per token.

OpenAI Developer Community; InfoQ, 2 Oct 2026
Fit

Which screen tasks suit a computer use agent

What a computer use agent is good for shows in the examples the vendors pick themselves, and those look strikingly alike. Microsoft lists automated data entry, invoice processing and data extraction, and supplies three sample instructions. Move invoice data from a PDF into a form. Add several items to an inventory system. Return two values from a table as plain text. These are not future visions but the cases the vendors demonstrate with (and you can tell, because all three are tame and forgiving).

That tells you what a suitable task looks like. It repeats, it always runs through the same screens, it has a checkable result, and a mistake would be cheap to correct. Conversely, anything that calls for case-by-case judgement, touches money or contracts, or cannot be withdrawn after submission is a poor fit. Microsoft adds that computer use works best for autonomous agents performing tasks in the background.

One practical note that follows from the same documentation: extraction is a good place to start. It changes nothing in the target system, the result can be held against the source, and a wrong turn costs only a second run. So start with reading, not with writing.

Type of workFitWhy
Read values from a portal and pass them onGood starting pointChanges nothing in the target system, result checkable against the source
Move data from a PDF into a formSuitableRepeats, fixed fields, errors visible before submitting
Maintain master data in an ageing ERP systemSuitable with approvalAlways the same screens, but write access to live data
Place orders or trigger paymentsOnly with case-by-case consentVendors explicitly require human confirmation here
Judgement calls such as assessing complaintsNot suitableNeeds judgement and accountability, not clicking
Permissions

What permissions a computer use agent actually needs

Permissions are the part of a computer use agent you cannot delegate to the technology. An agent that clicks acts with the permissions of the account it runs under, so it can do precisely as much damage as that account is allowed to do. All three vendors arrive at the same advice independently, which is itself worth noting. An isolated browser or VM plus an allow list of sites and actions, says OpenAI. A dedicated virtual machine or container with minimal privileges, says Anthropic. And in the Microsoft documentation it is dedicated, isolated machines with an account on the principle of least privilege.

Credentials are where it gets concrete. Microsoft offers two routes, an encrypted internal store and your own Azure Key Vault, and warns in the same place about a setting that is easy to miss: if the tool uses its maker's credentials and the agent is then shared, anyone using it can act with that maker's access on the configured machine. So before you make an agent available to a team, check whose account it actually runs under.

One limit of the allow list belongs on the table too, because it is easily misread. Microsoft writes that access control only prevents the model from taking actions on sites and apps that are not on the list, but it does not stop the model from opening them. If only one site is listed, the agent can still navigate elsewhere, and only the attempt to interact there fails. That is real protection, just not a fence around the machine.

Building blockWhat it achievesWhere it is documented
Dedicated, isolated machineA wrong turn stays in that environment and reaches no other systemsOpenAI, Anthropic and Microsoft
Low-privilege accountThe agent can only do what the task requiresAnthropic and Microsoft
Allow list of permitted sites and appsActions outside the list failOpenAI, Anthropic and Microsoft
Credentials in a secret storePasswords never appear in the instructions given to the agentMicrosoft, internal store or Azure Key Vault
Encrypted connections onlyThe agent does not work on unencrypted sitesMicrosoft, the Enforce HTTPS setting
Limits on steps, time and costA run that goes astray ends by itselfOpenAI
Approval

Where a person must approve before a computer use agent submits anything

On approvals the vendors of computer use agents are unusually direct. Consequential actions must be confirmed and users kept in control of purchases, data transmission and destructive changes, writes OpenAI, and the confirmation belongs at the point where the risk arises. Anthropic gets more specific and spells out what a human should agree to. Decisions that might result in meaningful real-world consequences, plus any task requiring affirmative consent, so accepting cookies, completing financial transactions or agreeing to terms of service.

Microsoft has built a mechanism for it. When the agent needs a confirmation or is missing a detail, it sends a review request by email to a named reviewer, carrying the agent name, a link to the run and a field for the answer. The workflow stays paused, and if the stated time limit passes without a response, the run ends. Worth reading too: turning this oversight off does not make the agent autonomous. It still pauses when it needs confirmation, and with no reviewer configured the session simply fails.

And now the sentence that matters. In the same documentation Microsoft writes that model behaviour is probabilistic, so these requests do not fire in every situation where a person would want a pause, and may fire when no pause is needed. It then warns plainly not to rely on human review or clarification requests as a safety measure, or as a guarantee that the system always asks for human input before proceeding. We think that is the most important line in the whole documentation set, because it moves the safeguard to where it is dependable: into permissions and the allow list. An agent that technically cannot submit an order needs no question about whether it should.

No guarantee

Microsoft explicitly warns against treating the built-in request for human input as a safety measure.

Microsoft Learn, Human supervision for computer use
Decision

When a computer use agent pays off for your business

Whether a computer use agent pays off rests on a single prior question, and it has nothing to do with AI. Is there an API? If yes, that is the quieter road, and the agent would be the detour through the interface. The technology gets interesting exactly where no API exists and probably none is coming, so with ageing line-of-business software, with supplier and government portals, and with programs maintained by a small vendor (where an interface is usually not even planned).

The second point is volume. Every run consists of many small steps, and those cost money. Microsoft bills five Copilot Credits per step and fifteen with a premium model, and gives a worked example: a timesheet run made up of opening the browser, creating a new sheet, filling the fields and submitting comes to four steps and therefore twenty credits. At Anthropic a single screenshot alone runs to roughly one thousand to one thousand eight hundred tokens. A task that comes up twice a month will not pay for itself. One that comes up forty times a day will.

  1. 01

    No API, task runs daily

    Worth examining

    This is where computer use earns its keep. Start with a read-only case, measure the hit rate for a week, and only then extend to steps that write.

  2. 02

    An API exists

    Use the API first

    Faster, cheaper and more dependable. Computer use remains the route for the leftovers the API does not cover.

  3. 03

    Task is rare and touches money or contracts

    Low priority

    The effort for machine, permissions and approvals stays the same while the benefit is small. These cases stay with people for now.

Limits

Where computer use agents still fall short, from prompt injection to Citrix

The first limit of a computer use agent is an attack route with a name of its own. Prompt injection means hidden instructions in a web page, an image or a screenshot steering the agent somewhere else. In some circumstances the model will follow commands found in content even when they conflict with your instructions, Anthropic describes, and therefore runs classifiers that scan what the tools return to flag potential injections and make the agent check back when something looks wrong. The same warning appears at Microsoft, there tied to a recommendation for trusted, isolated environments. OpenAI puts the underlying stance most briefly. Treat screen content as untrusted, because text in a page cannot grant permission.

The second limit is technical, and it bites mid-sized companies harder than expected. Microsoft writes that password fields are supported on all websites and most Windows applications, but names app types where that may not hold: Electron, Java, Unity, games, command-line interfaces, Citrix and other virtualised environments. If your line-of-business software is delivered over Citrix, check that before anything else. Anthropic adds that small UI elements are tricky to hit and that heavy downscaling of high-resolution displays hurts click precision.

The third limit is expectation. An agent driving an interface works more slowly than an API and makes mistakes an API would not, because it interprets a picture instead of reading a data field. So every deployment needs a check on the result, and OpenAI says as much: set step, time or cost limits, support cancellation, and check the actual outcome. Plan for that and you get a useful tool. Skip it and you get a source of errors nobody notices.

If you are wondering where your own company should start, the wider picture helps: which tasks belong with AI at all is sorted in AI in mid-sized companies, the line between goal-driven systems and simple helpers in agentic AI, and what an office assistant delivers today in AI assistant for business.

Frequently asked questions

Frequently asked questions

What does "computer use agent" mean?

A computer use agent is an AI system that operates software through its interface, clicking and typing with a virtual mouse and keyboard. It looks at a screenshot, picks the next action, and checks on the following screenshot what happened. It needs no API to do this.

How is computer use different from classic process automation?

Conventional automation follows a fixed path: this button at this position. Change the interface and the script breaks. A computer use agent interprets the picture instead, which is why Microsoft writes in its own documentation that the tool adapts to interface changes and keeps working when buttons or screens change. In exchange it is slower and its outcome less predictable.

Which Windows software can a computer use agent operate?

In principle any software a person can operate, so websites and Windows applications. Microsoft does name app types where password fields may not work, among them Electron, Java, Unity, command-line interfaces and Citrix as well as other virtualised environments. If your line-of-business software runs over Citrix, check that before you plan.

Can you automate software without an API?

That is exactly what a computer use agent is for. Because it works through mouse and keyboard, it needs no programming interface, only a surface a person could operate. Microsoft names programs with no direct connection as a core use case. Where an interface does exist, though, it stays the faster and more dependable route.

How to build a computer use agent?

Two routes are common. Through an API you run the loop yourself, sending the task, executing the returned actions and passing back a screenshot, with OpenAI naming gpt-6.1-sol and Anthropic shipping the computer_toolset_20260801 toolset. In Microsoft Copilot Studio you add computer use as a tool, describe the task in plain language and pick the model, with no code involved. Either way the machine, the permissions and the allow list are settled before the first run.

How much does a computer use agent cost?

It depends on the route. Microsoft bills in Copilot Credits, five per step and fifteen with a premium model, and cites a timesheet example of four steps and therefore twenty credits. Through the OpenAI API you pay for tokens, with GPT-6.1 Sol reported at DevDay to sit at one fifth of GPT-6 Astra's standard prices. Cost is driven by the number of screenshots, not by the individual click.

Can a computer use agent cause damage?

It can do whatever its account is allowed to do. That is why OpenAI, Microsoft and Anthropic all require a dedicated, isolated machine, a low-privilege account and an allow list of permitted sites and applications. The safeguard lives in those three things, not in the model's good intentions.

Does the agent need my staff's passwords?

No, and it should not receive them in the instructions either. Microsoft stores credentials in an encrypted internal store or in an Azure Key Vault and enters them itself when a sign-in prompt appears. Anthropic goes further and advises avoiding giving the model access to sensitive data such as account login information at all.

Will the agent ask a human before doing something important?

You cannot count on it. Microsoft writes that model behaviour is probabilistic, so requests do not come in every situation and sometimes come without need, and warns explicitly against treating those requests as a safety measure. Whatever must not happen is prevented through permissions, not through a question.

Is a desktop AI agent the same thing?

Broadly yes, when the term means software that drives a desktop with a virtual mouse and keyboard. Anthropic draws one distinction worth knowing: for work that stays inside web pages, a browser tool is the closer fit, because it reads and acts on the page itself and needs no full desktop environment.

Is an AI agent with computer use ready for production?

For reading tasks on a dedicated machine, many teams will find it usable today. For anything that writes to live data or submits something, the vendors' own documentation is the better guide than any benchmark. Isolate the machine, keep permissions low, use an allow list, and have a person approve what cannot be undone.

Next steps

How to test a computer use agent on a first task

  1. 01

    Rule out an API

    Ask the vendor of the software in question whether an API exists. If one does, it is the faster and more dependable route, and the computer use question answers itself.

  2. 02

    Pick a read-only task

    Find a recurring job that only reads data and changes nothing. The result can be held against the source, and a wrong turn costs only a second run.

  3. 03

    Settle machine and permissions first

    Set up a dedicated, isolated machine, give the account only the permissions needed, put the permitted sites and apps on the allow list, and keep credentials in a secret store.

  4. 04

    Measure the hit rate for a week

    Let the agent run alongside while a person still does the same work, then compare results. Only take over once the rate holds, and keep approval on every step that writes.

An agent that clicks opens up the software that never got an API. What it can do in there is something you decide beforehand through permissions, not afterwards through a question.

or call: +49 151 1576 5566

Further reading

Projekt-Detail

    Got a project in mind?

    We reply personally. First a use-case check, then an architecture proposal.

    Start your inquiry