# Mantaq > Mantaq builds document and workflow automation systems for teams that run on paperwork. Documents come in, a person reviews them, and they go out finished, with every step recorded. Mantaq is a small software team based in Islamabad, Pakistan, building for clients in the US, UK and EU. Every system has three parts: document automation, the platform it runs on, and the interface the team uses every day. Clients own the code, delivered to their own accounts and infrastructure. ## Pages - [Home](https://www.mantaq.co/): What Mantaq builds and who it is for - [Services](https://www.mantaq.co/services): Document automation, platform development and interface design, plus common questions - [About Us](https://www.mantaq.co/about-us): The team and how it works - [Blog](https://www.mantaq.co/blogs): Articles from the Mantaq team ## Case studies - [PeopleWorks](https://www.mantaq.co/case-studies/peopleworks): HR letter generation, signing and reporting for dozens of client companies, cut from 20 minutes per letter to 20 seconds - [VA Claims Made Easy](https://www.mantaq.co/case-studies/va-claims-made-easy): HIPAA-compliant AI platform that simplifies VA disability claims for veterans - [Compliance Portal](https://www.mantaq.co/case-studies/compliance-portal): Credential, signature and audit tracking across three behavioral health companies - [Stay Funded 360](https://www.mantaq.co/case-studies/stay-funded-360): Receipt capture and monthly grant reconciliation packets, built for a grant-funded nonprofit and now a product ## Contact - Email: contact@mantaq.co - Book a 30-minute call: https://www.mantaq.co/#contact ## Articles in full --- title: "GPT-6 Astra can use a computer. What that changes for your business" description: "OpenAI's GPT-6 Astra is built to do work on a computer, not just answer questions. This post covers what it can do, what it costs, where the limits are, and how to test it on real work." author: "Usman Masood, Developer" published: 2026-10-02 updated: 2026-10-02 url: https://www.mantaq.co/blogs/gpt-6-astra-for-business --- # GPT-6 Astra can use a computer. What that changes for your business OpenAI's GPT-6 Astra is built to do work on a computer, not just answer questions. This post covers what it can do, what it costs, where the limits are, and how to test it on real work. OpenAI released GPT-6 Astra to the public on September 4, 2026. It is a model built to finish tasks on a computer: open apps, click through screens, fill in forms and run multi-step workflows. This post covers what Astra does, what it costs, where the risks are, and how a business can test it without betting a process on it. ## Astra is built to do work, not to chat OpenAI sums up the model in one line: "Anything you can do on a computer, Astra can do for you." Earlier models mostly answered questions or wrote text you then had to use yourself. Astra operates software through the screen, the same way a person does. That matters because most business software has no clean API. Old accounting tools, government portals and internal admin panels only have a screen. A model that can read that screen and click the right button can work with them without anyone building an integration first. OpenAI says Astra is better than its earlier models at: - **Staying on task** across long workflows with many steps. - **Respecting task boundaries**, so it does what was asked and stops. - **Understanding intent** when the request is loose or incomplete. - **Handling tedious work** such as data entry and form filling. Astra comes as one model for now. OpenAI released two more GPT-6 models, Sol and Luna, on September 22. ## The benchmark numbers come from OpenAI The headline result is on OSWorld 2.0, a test of real computer tasks. OpenAI reports Astra at 72.6 percent, up from 65.7 percent for GPT-5.6 Sol, its previous top model. On GPQA Diamond, a set of graduate-level science questions, OpenAI reports 96.0 percent. Read these numbers with care. OpenAI ran these tests itself, and independent results take time to appear. Some of the most quoted figures also depend on the setup. One example is 99.9 percent on ARC-AGI-3, a puzzle-solving test. It used a special test harness that keeps state between steps and is expensive to run. A normal API call will not get that score. Greg Brockman, OpenAI's president, said Astra could eventually be seen as the arrival of artificial general intelligence. That is a claim about the future, not a measured result. For a business, a better test is simpler: can it finish your tasks, correctly, more often than it fails? ## Speed is the part your team will notice OpenAI says Astra takes about 40 minutes for an OSWorld 2.0 task that took GPT-5.6 Sol about 75 minutes. That is roughly 47 percent less time per task. A faster agent changes how people use it. A task that takes over an hour gets started and forgotten. A task that takes 40 minutes can run while someone finishes a meeting and checks the result straight after. Shorter runs also mean less time for an agent to drift off course, which is where many agent failures happen. Reported results also say Astra uses far fewer output tokens than some competing models on the same agent tasks. Fewer tokens means a lower bill for the same work. ## The cost adds up on long tasks Astra costs $10 per million input tokens and $50 per million output tokens through the OpenAI API. Cached input costs $1 per million. It can read up to about 1.05 million tokens at once and write up to 128,000 tokens in one reply. An agent working through screens reads a lot. Every screenshot and every page it looks at counts as input. Take a long task that reads 2 million tokens and writes 200,000. It costs about $30 at standard prices. That is cheap compared with an hour of a person's time. It is expensive if the task fails halfway and has to run again. Measure the cost of a finished task, not the price per token. A cheaper model that needs three attempts can cost more than Astra getting it right the first time. Astra is also in ChatGPT on the Plus, Pro, Business and Enterprise plans, and on Amazon's cloud. ## Strong models come with tighter limits Astra is the first OpenAI model rated "Critical" on the company's own cybersecurity scale. A model that can find and use security flaws well can help defenders and attackers alike. OpenAI has limited access to those abilities and the model refuses some security requests. AI safety researchers have raised a second concern. Astra uses a new reasoning method that makes its step-by-step thinking harder to inspect. When a model acts on your systems, being able to see why it took an action matters for finding and fixing mistakes. For a business, this means three practical rules: - **Give the agent its own account** with only the permissions the task needs. - **Log every action** it takes so a person can review the trail. - **Keep a person on the final step** for anything that sends money, emails customers or deletes data. ## Start with one dull workflow Pick a task your team does every week, that follows the same steps each time, and that lives in software with no API. Good candidates include copying data between two systems, filling in the same portal forms, or pulling reports from several dashboards into one sheet. Run Astra on that one task for two weeks with a person checking every result. Track three numbers: how often it finishes correctly, how long it takes, and what each finished run costs. Those numbers tell you whether to expand, adjust or stop. Mantaq builds AI workflows where the model does the repetitive steps and people approve the results. If you have a process in mind and want to know whether Astra or another model can run it, [get in touch with us](https://www.mantaq.co). --- --- title: "AI is changing work faster than companies are changing with it" description: "Almost every company now uses AI, but few have changed how their work runs. The latest research on adoption, jobs and productivity, and what the teams getting results do differently." author: "Usman Masood, Developer" published: 2026-10-01 updated: 2026-10-01 url: https://www.mantaq.co/blogs/ai-is-changing-work --- # AI is changing work faster than companies are changing with it Almost every company now uses AI, but few have changed how their work runs. The latest research on adoption, jobs and productivity, and what the teams getting results do differently. Almost every large company now uses AI somewhere. Far fewer have changed how their work actually runs, and that gap decides who gets value from it. This post brings together the latest public research on adoption, jobs and productivity, and ends with the steps that move a team from trying AI to running on it. ## Nearly every company uses AI, few see results AI in business went from experiment to default in about three years. McKinsey's [State of AI 2025 survey](https://www.mckinsey.com/capabilities/operations/our-insights/the-state-of-ai) found that 88 percent of organizations use AI in at least one business function, up from 78 percent a year earlier. Stanford's [2026 AI Index](https://hai.stanford.edu/ai-index/2026-ai-index-report) reports the same 88 percent. Results lag far behind. In the same McKinsey survey, only 39 percent of companies could point to any effect on their bottom line. Most are still running pilots: a chatbot in one team, a writing assistant in another, nothing connected to the core of the business. People outside work moved just as fast. Stanford reports that generative AI reached 53 percent of the population within three years of launch. ## The technology is no longer the bottleneck The models improved sharply in a single year. Stanford's index tracks two tests that map closely to real work: - **Coding:** on SWE-bench Verified, where models fix real software bugs, scores rose from 60 percent to near 100 percent. - **Computer tasks:** on OSWorld, where an AI agent completes tasks inside real desktop apps, success went from 12 percent to about 66 percent. Money followed the progress. US private investment in AI reached $285.9 billion in 2025, more than 23 times China's $12.4 billion, according to the same report. A company stuck in pilots today is rarely limited by what the tools can do. ## Jobs are shifting more than they are disappearing The World Economic Forum's [Future of Jobs Report 2025](https://www.weforum.org/press/2025/01/future-of-jobs-report-2025-78-million-new-job-opportunities-by-2030-but-urgent-upskilling-needed-to-prepare-workforces/) expects 170 million new roles and 92 million displaced roles by 2030, a net gain of 78 million. That churn touches 22 percent of today's jobs. The IMF [estimated](https://www.cnbc.com/2024/01/15/imf-warns-ai-to-hit-almost-40percent-of-global-employment-worsen-inequality.html) that almost 40 percent of jobs worldwide are exposed to AI. In advanced economies the figure is about 60 percent, against 26 percent in low-income countries. Exposed does not mean replaced. The IMF expects roughly half of the exposed jobs in advanced economies to gain from AI help, while the other half face lower demand. The pressure lands hardest at the entry level. Stanford researchers found in [Canaries in the Coal Mine](https://digitaleconomy.stanford.edu/publication/canaries-in-the-coal-mine/) that workers aged 22 to 25 in the most AI-exposed jobs, such as software development and customer service, saw a 13 percent relative drop in employment. Older workers in the same jobs held steady or grew. Skills move with the jobs. The WEF expects nearly 40 percent of the skills used at work today to change by 2030, and 63 percent of employers already name the skills gap as their biggest barrier to change. ## AI helps the least experienced people most When AI sits inside the daily workflow, the gains are measurable. Erik Brynjolfsson, Danielle Li and Lindsey Raymond studied 5,179 customer support agents in [Generative AI at Work](https://www.nber.org/system/files/working_papers/w31161/w31161.pdf). An AI assistant raised the number of issues each agent resolved per hour by 14 percent on average. New and less skilled agents improved by 34 percent. Experienced agents barely changed. The assistant had learned from the best agents' conversations, so it passed their habits on to everyone else. The study also found happier customers and fewer agents quitting. This changes how to read the entry-level numbers above. AI is taking over some starting tasks, and it also brings new people up to speed much faster. Companies that use it to train junior staff, not only to replace them, keep their talent pipeline working. ## Teams that get results change the work, not just the tools The gap between using AI and getting value from it comes down to a few choices: - **Redesign the workflow.** Adding a chatbot to an old process gives small gains. Rebuilding the process so AI handles the repetitive steps gives large ones. - **Start with one costly, repetitive process.** Pick work that eats hours every week, such as reading forms, matching receipts or answering the same questions. Measure how long it takes before you change anything. - **Keep people on the decisions.** Let AI draft, sort and suggest. Let a person approve anything that affects money, customers or compliance. Stanford counted 362 documented AI incidents in 2025, up from 233 the year before. - **Train the team as you roll it out.** Employers rank the skills gap as their top barrier, and it does not close by itself. - **Measure in hours and errors.** A result is a number, such as hours saved per week or mistakes caught per month. ## Pick one process and rebuild it Most companies already pay for AI tools. The next step is choosing one process and redesigning it so the tool does the repetitive part and people own the judgment calls. Mantaq builds AI automation, internal tools and integrations for teams at exactly this point. If a process in your company eats hours every week, [talk to us](https://www.mantaq.co/#contact) and we will map where AI fits. --- --- title: "Why we taught our AI receipt reader to say \"I can't read this\"" description: "A blurry receipt photo made our AI reader return a confident wrong total. We found the cause, changed three things, and learned why \"I can't read this\" is a better answer than a guess." author: "Usman Masood, Developer" published: 2026-10-01 updated: 2026-10-01 url: https://www.mantaq.co/blogs/ai-receipt-reader-says-i-cant-read-this --- # Why we taught our AI receipt reader to say "I can't read this" A blurry receipt photo made our AI reader return a confident wrong total. We found the cause, changed three things, and learned why "I can't read this" is a better answer than a guess. We build a grant reconciliation app for nonprofits. Staff log their expenses, attach receipts, and at the end of each month the app puts together a packet for the funder. To save typing, the app reads each receipt with AI and suggests the subtotal, tax and total. Most of the time this works well. But receipts in real life are not clean scans. They are phone photos taken in a dim restaurant, folded in a wallet for a week, or shrunk down by a chat app. So we decided to test the reader on the kind of receipts people actually upload. ## The problem: confident and wrong One test receipt had a total of $204.75. We made a low resolution copy of it, the kind of photo you get when an image is shrunk and then blown up again. The reader looked at it and suggested a total of $22.50. There was no warning and no sign of doubt. It just gave a clean, believable number that was off by about $180. This is the worst kind of mistake an AI feature can make. If the reader says nothing, the person types the amount in themselves. If it says something wrong and sounds sure, the person may accept it, and the wrong number ends up in a report that goes to a funder. So the goal became clear. When the reader can read a receipt, it should get it right. When it can't, it should say so. ## How we tested it We made a small set of test receipts from one clear original. There was the clear version, a lightly blurred one, a medium blur, a heavy blur and a low resolution copy. We also made a receipt with the totals cropped off and a file with two receipts in it. Then we ran each one through the real reader several times. Running the same file more than once matters. An AI model does not always give the same answer twice, and a feature that works one time in three is not a feature people can trust. The results were mixed. Medium and heavy blur were handled well, since the reader declined to give an amount. The clear receipt was mostly right. But the light blur and the low resolution copies were unpredictable. Sometimes the read failed, sometimes it gave a close answer, and sometimes it gave a confident wrong one. ## The hidden cause: the AI ran out of room When we looked at the failed reads, the cause was not what we expected. It was not mainly the blur. The model we use thinks through a problem before it answers, and that thinking counts toward a limit on how much it can write back. Our limit was 600 tokens. On a hard photo, the model spent all 600 working out what it was looking at and ran out of room before it could give an answer. About half of the photo reads were being cut off this way. It also explained why the same photo could pass on one run and fail on the next. Some runs needed a little more thinking than others, and those were the ones that hit the limit. ## Three fixes We made three changes to the reader. **More room to think.** We raised the limit from 600 to 2000 tokens. The model now has space to work through a hard photo and still give a full answer. **Full detail on photos.** The reader now sends photos at full detail. Before this, a photo could be scaled down before the model ever saw it, which threw away the small print that matters most on a receipt. **A clear rule to decline.** We told the reader directly that if it can't read the amounts with confidence, it must say so and return nothing. A guess is not allowed. After these changes we ran the same tests again. The low resolution copy that had produced $22.50 now declined on every run. The clear receipt gave the correct amounts on every run. Medium and heavy blur still declined, which is what we want. ## Cleaning up blurry photos first That left the lightly blurred photos. These are the most common problem in practice, and they are often readable by a person, so declining every time felt like giving up too early. We tried cleaning the photo before sending it. We tested a few versions on a lightly blurred photo, ten runs each. Enlarging the photo on its own gave the right amounts once in ten runs. Enlarging it, boosting the contrast and then smoothing out the noise gave the right amounts five times in ten. Most of the other runs declined, which is the safe result. Some ideas made things worse. Boosting contrast without smoothing afterwards caused every read of the blurry photo to fail. Sharpening the photo did the same. The lesson for us was to test each step on its own and not assume that a clearer looking photo is easier for the AI to read. Cleanup does have a ceiling. Medium blur and very low resolution photos still could not be read, whatever we did to them. For those the reader declines, and the person types the amount in. ## What we accepted, and why it's fine The reader is not perfect, and we wrote down exactly where it still falls short. A lightly blurred photo can still be off by a few cents. A receipt with a fancy script logo can still produce the wrong shop name. Heavily blurred and tiny photos are not read at all. We are comfortable with this because of one rule the whole app follows. The AI suggests, the person confirms, and nothing saves on its own. Every amount the reader finds is shown as a suggestion that the person checks before applying. A few cents off is easy to spot when you are looking at the receipt. A confident $22.50 on a $204.75 receipt is much easier to miss, and that is the mistake we removed. We also dropped a feature we had planned. We thought we might need a way to pick and add up single items on long receipts. Testing showed that real receipts print their own totals, and the reader already reads those. Not building it kept the app simpler. If you are adding AI to your own product, our advice is short. Test it on the messy inputs your users really have, run each test more than once, and make sure the AI is allowed to say "I don't know." If you want help doing that for your app, [talk to us at Mantaq](https://www.mantaq.co). --- --- title: "How AI Helps a Business Grow Without Hiring More People" description: "The newest AI models are faster and cheaper than ever. Here is what that means for a growing business, with a real example from a platform we built." author: "Usman Masood, Developer" published: 2026-10-01 updated: 2026-10-01 url: https://www.mantaq.co/blogs/ai-models-for-scaling-business --- # How AI Helps a Business Grow Without Hiring More People The newest AI models are faster and cheaper than ever. Here is what that means for a growing business, with a real example from a platform we built. September was a busy month for AI. OpenAI, Anthropic and Google all released new models within a few weeks of each other. Each one claims to be faster, smarter or cheaper than the last. If you run a business, the release notes are not the useful part. The useful question is simpler. Can this help my team take on more work without hiring more people? The short answer is yes, if you use it on the right work. Here is what came out, where it actually helps, and what it looked like on a real project. ## The new models, in plain words Here are the main releases from the last few weeks: - **GPT-6 Sol and GPT-6 Luna** from OpenAI came out on September 22. Sol is for everyday professional work and coding. Luna is the small, low-cost option. OpenAI followed up with GPT-6.1 Sol a week later. - **Claude Opus 5.5** from Anthropic came out on September 22. It is built for long, complex work like coding and research. Anthropic says it costs about 40% less to run than the model before it on typical work. - **Claude Sonnet 5.5** followed on September 28. It is the faster, cheaper option for well-defined everyday tasks like documents and spreadsheets. - **Gemini 3.8 Flash** from Google came out earlier in September. It is Google's fast, low-cost model for high-volume work. - **Gemini 4 Argon** was announced by Google on September 30. For now it is only open to a small group of security partners. The pattern matters more than any single name. Every company now ships a big model for hard thinking and a smaller one for fast, cheap work. And prices keep coming down. OpenAI cut API prices for Sol and Luna in half compared with the promotional pricing of the generation before. ## Scaling used to mean hiring For most businesses, growth has followed one rule. More customers means more work, and more work means more people. That works until it doesn't. Each new hire takes time to find and train. Costs grow as fast as revenue. And a lot of what new people end up doing is the same task, over and over. Copying data from one form into another. Reading a document to find three facts. Writing the same kind of email for the hundredth time. That repetitive part is where AI fits. It does the reading, sorting and first drafts. Your team does the checking, the judgment and the work with customers. The team stays the same size while the amount of work it can handle grows. ## Where AI actually helps AI is strongest at work that is repetitive and has a clear right answer. Some good places to start: - **Reading documents.** Pulling names, dates, amounts or conditions out of forms, PDFs and scans. - **First drafts.** Letters, summaries, reports and replies that a person then checks and sends. - **Sorting requests.** Reading incoming emails or tickets and sending each one to the right person. - **Answering common questions.** A chatbot that handles the questions your team answers every day, and hands the rest to a person. What these have in common is volume. If a task happens five times a month, automating it saves little. If it happens five hundred times, it changes how big your team needs to be. ## A real example: VA Claims Made Easy [VA Claims Made Easy](https://www.mantaq.co/case-studies/va-claims-made-easy) helps veterans file disability claims. Many veterans give up on benefits they have earned because the paperwork is long and confusing. The founders wanted a platform that made the process something a veteran could actually finish. The slow part was the medical records. Every claim starts with documents, and someone has to read them to find the conditions that matter. Done by hand, that is careful data entry for every single veteran. We built a system where AWS Textract reads the uploaded records and pulls out the conditions automatically. The platform then asks the veteran follow-up questions based on those conditions, and the statements and supporting documents are drafted automatically. Then a person steps in. A VA agent or doctor reviews and edits every document before the veteran sees it. A wrong claim can cost a veteran their benefit, so the AI never has the final word. The result was 80% faster document handling compared with manual data entry, and the full platform launched in two months. The agents and doctors spend less time typing and more time on the review only they can do. That is what scaling with AI looks like. ## Picking a model is the easy part With so many releases, it is easy to think the choice of model is the big decision. Usually it is not. Most of the current models are good enough for reading documents and writing drafts, and switching from one to another later is often a small change. The parts that decide whether a project works sit around the model: - **The workflow.** Which step does the AI do, and what happens before and after it? - **The check.** Who reviews the output, and how easy is it for them to fix a mistake? - **The data.** Is the information the AI reads clean, complete and safe to send? For VA Claims Made Easy, the records were medical data, so security had to be part of the design from day one. Get these right and you can swap in a newer model whenever one comes out. Get them wrong and the best model in the world will not save the project. ## Where to start You do not need a big AI strategy to begin. Start small: 1. **Pick one repetitive task.** Something your team does many times a week that follows the same steps. 2. **Measure it.** How long does it take now, and how often does it happen? 3. **Automate the boring part.** Let AI do the reading or drafting, and keep a person on the final check. 4. **Compare the numbers.** If it saves real time, move on to the next task. If you want help finding that first task, [book a meeting with us](https://www.mantaq.co/#contact). We will look at how your team works today and tell you honestly where AI would help and where it would not.