A laptop on a desk showing lines of code on the screen
Article

GPT-6 Astra can use a computer. What that changes for your business

OpenAI's GPT-6 Astra is built to do work on a computer, not just answer questions. This post covers what it can do, what it costs, where the limits are, and how to test it on real work.

Usman MasoodDeveloper
View as Markdown

Astra is built to work inside software, through the screen.

OpenAI released GPT-6 Astra to the public on September 4, 2026. It is a model built to finish tasks on a computer: open apps, click through screens, fill in forms and run multi-step workflows. This post covers what Astra does, what it costs, where the risks are, and how a business can test it without betting a process on it.

Astra is built to do work, not to chat

OpenAI sums up the model in one line: "Anything you can do on a computer, Astra can do for you." Earlier models mostly answered questions or wrote text you then had to use yourself. Astra operates software through the screen, the same way a person does.

That matters because most business software has no clean API. Old accounting tools, government portals and internal admin panels only have a screen. A model that can read that screen and click the right button can work with them without anyone building an integration first.

OpenAI says Astra is better than its earlier models at:

  • Staying on task across long workflows with many steps.

  • Respecting task boundaries, so it does what was asked and stops.

  • Understanding intent when the request is loose or incomplete.

  • Handling tedious work such as data entry and form filling.

Astra comes as one model for now. OpenAI released two more GPT-6 models, Sol and Luna, on September 22.

The benchmark numbers come from OpenAI

The headline result is on OSWorld 2.0, a test of real computer tasks. OpenAI reports Astra at 72.6 percent, up from 65.7 percent for GPT-5.6 Sol, its previous top model. On GPQA Diamond, a set of graduate-level science questions, OpenAI reports 96.0 percent.

Read these numbers with care. OpenAI ran these tests itself, and independent results take time to appear. Some of the most quoted figures also depend on the setup. One example is 99.9 percent on ARC-AGI-3, a puzzle-solving test. It used a special test harness that keeps state between steps and is expensive to run. A normal API call will not get that score.

Greg Brockman, OpenAI's president, said Astra could eventually be seen as the arrival of artificial general intelligence. That is a claim about the future, not a measured result. For a business, a better test is simpler: can it finish your tasks, correctly, more often than it fails?

Speed is the part your team will notice

OpenAI says Astra takes about 40 minutes for an OSWorld 2.0 task that took GPT-5.6 Sol about 75 minutes. That is roughly 47 percent less time per task.

A faster agent changes how people use it. A task that takes over an hour gets started and forgotten. A task that takes 40 minutes can run while someone finishes a meeting and checks the result straight after. Shorter runs also mean less time for an agent to drift off course, which is where many agent failures happen.

Reported results also say Astra uses far fewer output tokens than some competing models on the same agent tasks. Fewer tokens means a lower bill for the same work.

The cost adds up on long tasks

Astra costs $10 per million input tokens and $50 per million output tokens through the OpenAI API. Cached input costs $1 per million. It can read up to about 1.05 million tokens at once and write up to 128,000 tokens in one reply.

An agent working through screens reads a lot. Every screenshot and every page it looks at counts as input. Take a long task that reads 2 million tokens and writes 200,000. It costs about $30 at standard prices. That is cheap compared with an hour of a person's time. It is expensive if the task fails halfway and has to run again.

Measure the cost of a finished task, not the price per token. A cheaper model that needs three attempts can cost more than Astra getting it right the first time.

Astra is also in ChatGPT on the Plus, Pro, Business and Enterprise plans, and on Amazon's cloud.

Strong models come with tighter limits

Astra is the first OpenAI model rated "Critical" on the company's own cybersecurity scale. A model that can find and use security flaws well can help defenders and attackers alike. OpenAI has limited access to those abilities and the model refuses some security requests.

AI safety researchers have raised a second concern. Astra uses a new reasoning method that makes its step-by-step thinking harder to inspect. When a model acts on your systems, being able to see why it took an action matters for finding and fixing mistakes.

For a business, this means three practical rules:

  • Give the agent its own account with only the permissions the task needs.

  • Log every action it takes so a person can review the trail.

  • Keep a person on the final step for anything that sends money, emails customers or deletes data.

Start with one dull workflow

Pick a task your team does every week, that follows the same steps each time, and that lives in software with no API. Good candidates include copying data between two systems, filling in the same portal forms, or pulling reports from several dashboards into one sheet.

A man working at a computer in an office
Keep a person checking every result during the trial.

Run Astra on that one task for two weeks with a person checking every result. Track three numbers: how often it finishes correctly, how long it takes, and what each finished run costs. Those numbers tell you whether to expand, adjust or stop.

Mantaq builds AI workflows where the model does the repetitive steps and people approve the results. If you have a process in mind and want to know whether Astra or another model can run it, get in touch with us.