Claude Fable 5.1 and GPT-6 Astra shipped two days apart on Sept. 1 and 3, 2026. The two have identical API prices, and measuring against benchmarks that matter in the tax and accounting world, both products win and lose in certain categories. The decision that matters the most is what either model is allowed to do inside a firm’s workflow, and guardrails always must be considered.
Key takeaways
- Identical API list prices: $10 per million input tokens, $50 per million output; Claude’s cache reads cost 75% less than OpenAI’s
- Fable 5.1 knows tax developments through June 2026; Astra’s knowledge stops April 30, 2026
- Astra scores higher on tax-return calculation, while Fable 5.1 scores higher on desk work that leads to deliverables
- Both vendors report small but nonzero rates of agents acting beyond their instructions; Hugging Face incident in July shows serious risks involved
- With Truss’ Max AI, firms are protected against some of the more concerning security risks of using either model, while using both models indirectly in areas where they’re the strongest
Astra’s release as a brand-new flagship model is making waves in a community that expands beyond just accounting, but Anthropic’s flagship model got a facelift as well. It’s all happening smack-dab in the middle of extension season, so the question of switching models is naturally creating a mixture of excitement and apprehension in Slack and Teams channels across firms nationwide. ChatGPT is still the most used chatbot on the whole, while Claude has the deeper footprint inside the largest accounting firms in the country. Deloitte, PwC, and KPMG are all Claude users.
Also worth noting: Thomson Reuters’ 2026 State of Tax Professionals report asserted that practitioners “may not be so interested in public chatbots such as Claude and ChatGPT” and care more about the AI native to the tools they already use. There’s a pervasive skepticism involved in the accounting profession that is key to unravel when considering technology changes.
How do Claude Fable 5.1 and GPT-6 Astra stack up on industry benchmarks?
GPT-6 Astra scores higher on tax calculation, with a 60% strictly correct return score to Fable 5.1’s 46% according to TaxCalcBench v2. Astra also has the edge on benchmarks that test whether models can finish multistep tasks in software, like Terminal-Bench and AutomationBench. Fable 5.1 leads Artificial Analysis’ general reasoning index 66 to 61, and GDPval for deliverables like memos and analysis 1764 to 1580. Neither model should prepare a full return unsupervised.
Where Astra leads
Where Fable 5.1 leads
“Models can’t calculate tax returns reliably today.”
Does knowledge cutoff date matter in tax work?
The short answer is yes, at least when the model is answering something from memory. Fable 5.1’s current knowledge cutoff (as of publication) is June 2026, while Astra’s is April 30, 2026. In between those dates, the JCT Blue Book on the One Big Beautiful Bill Act was released, while the IRS dropped its first guidance on AI use for practitioners. However, if web search is on, the knowledge gap closes significantly.
- May 28, 2026: JCT Blue Book (JCS-1-26), the 341-page general explanation of Public Law 119-21. Fable 5.1 trained on it; Astra did not.
- June 17, 2026: Oregon Environmental Council v. IRS vacated Notice 2025-42 on the wind and solar construction start-date safe harbor.
- June 24, 2026: OPR Alert 2026-19, the IRS Office of Professional Responsibility’s Introductory Guidelines for Responsible AI Use in Federal Tax Practice. We covered it in detail when it came out.
These gaps are not permanent, as OpenAI’s next snapshot will go past June 2026. Both models know about the original One Big Beautiful Bill Act, signed into law on July 4, 2025. Both vendors ship a web search feature, as Truss does in Max AI’s Research Mode. The cutoff is most impactful for a casual search function, which is, to be fair, the most common use in many firms. Some might open a chat window and ask a question without ever turning search on.
Which costs more, Claude Fable 5.1 or GPT-6 Astra?
Going purely off sticker price, neither immediately reads as more expensive. Both charge $10 per million input tokens and $50 per million output tokens. Seats start at $20 per month on both. Cache reads for Claude are $0.25 per million to OpenAI’s $1.00, but Astra was measured to be significantly cheaper per completed task by Artificial Analysis, $3.26 to $7.63.
ChatGPT Plus costs $20 per month and includes Astra in full, with Pro plans starting at $100. Claude Pro is also $20 per month, but gates Fable 5.1 behind usage credits. Max starts at $100 and includes Fable but caps it at half of the regular weekly usage limit. Astra seems to have a sizable advantage on cost per completed task, which could become more important as firms adopt agentic workflows.
Asking either model for a research memo would only cost a few cents, but the real cost comes in the form of time spent by a preparer when checking the answers. In June, the IRS made it clear that such a review was required, not optional.
Why is alignment important to consider?
Alignment refers to whether the model does only what you asked and stops there. Both vendors ran a common test: Gray Swan’s prompt-injection benchmark, which measures how often hidden instructions in a document or web page are able to push agents into harmful actions within 15 attempts. Fable 5.1 failed 1.0% of the time and GPT-6 Astra 8.5%, both in a sample size of just over 1,800 prompt injections.
Here’s the disaster scenario: A member of the firm asks an agent to perform some research on a client. In pursuit of the information it thinks is necessary, the agent reaches out directly to the client via tool call to satisfy its request. It’s technically following instructions, but in doing something that wasn’t explicitly ruled out, it diminishes the trustworthiness of the firm’s letterhead in addition to causing headaches.
Misalignment seems to be getting more serious as models get more capable. There have been a number of recent incidents in which the agent, in pursuit of its goals, decided the best path was to hack into another company to get the answers it was looking for. OpenAI’s September 3 system card explicitly called out these occurrences.
“In coding contexts, misalignment generally stems from a mix of overeagerness to complete the task and interpreting user instructions too permissively”
OpenAI’s own testing confirmed that in simulations built to trip up the model, GPT-6 Astra sent emails or messages it shouldn’t have been authorized to send in 1.4% of tasks, even with its confirmation policy on. That policy asks the model to double-check before “sending certain communications or making purchases.” Needless to say, 1.4% of 5,000 client emails in a season being unauthorized or out of order would create a ton of unnecessary strife.
The one common test we could find that both companies ran was the Gray Swan prompt-injection benchmark, measuring how often hidden instructions in documents and web pages were able to push agents into harmful actions within 15 tries. Fable 5.1 failed 1.0% of the time and Astra 8.5%, both with a sample of just over 1,800 trials. Anthropic’s own reporting concedes that Fable sometimes works around permissions checks, and the July Hugging Face incident, where more than 1,200 agents within an unreleased OpenAI model broke out of their test environment and hacked into another AI company’s computer systems, showed the danger of what can happen when agents are left unchecked. The fix, as ever, is that the model needs to ask a person before taking action.
“The question isn’t whether, if you use Astra or Fable, it’s going to turn into the Terminator on you. The question is, if you ask a legitimate question, to what lengths will it go to try to answer that question, including doing things that you might not necessarily want it to do?”
How does Truss prevent AI from acting without firm approval?
Truss treats any AI output that is intended to reach a client as a human’s decision, and proxies all agent actions through proprietary gateways that restrict what the agent can do. Max, which turns a response to a prompt into an email with one click, is scoped to a single client or entity. A person then has to go through the physical act of sending the message. Client data is never used to train models.
Truss is not contracted to a single enterprise AI model, and accounting firms shouldn’t have to re-evaluate their safety postures every time a vendor ships a new version of the model. That’s why Truss created guardrails at the product level that hold no matter which model comprises the engine.
- Drafts, but never sends. Max searches a client’s recent communications with the firm, shows its sources, and turns answers into ready-to-edit emails inside the client’s project. But while Max can read and search, it will never be able to send; only a person will be able to click that button.
- One client per request. Each request will only ever pull from the client whose project you are working in, which is the exact segmentation the IRS demanded from vendors in June 2026.
- Review step added to prep. Any return routed to an AI prep vendor through Truss flips to “Action Required” or “Review Needed” and notifies the firm to make sure a preparer’s judgment is always involved in the process.
- Other critical benchmarks. Client data is never used to train models, processing is always done on U.S. based infrastructure, and controls have been SOC 2, Type II audited since 2024.
What should determine whether a firm switches from Claude to GPT-6 Astra?
Astra could be the more appealing option for firms whose calculations run through the API, as Astra’s 60% TaxCalcBench score and lower cost per task are differentiators. Claude might be better for firms that mostly use AI for research and drafting, as Fable’s 1.0% injection attack rate is significantly lower. For most firms, the most important action to take is to iron out its approval step, regardless of which model it’s using.
| Factor | Claude Fable 5.1 | GPT-6 Astra | Edge |
|---|---|---|---|
| Release and knowledge | |||
| Release | September 1, 2026 | September 3 to 4, 2026 | Even |
| Knowledge cutoff | June 2026 | April 30, 2026 | Fable |
| Cost | |||
| API price | $10 in, $50 out per million tokens; $0.25 cache read | $10 in, $50 out per million tokens; $1.00 cache read | Fable on caching |
| Seat price | Pro $20 a month; Max from $100 | Plus $20 a month; Pro from $100 | Even |
| Cost per Artificial Analysis task | $7.63 | $3.26 | Astra |
| Calculations | |||
| TaxCalcBench v2, strict | 46% | 60% | Astra |
| Terminal-Bench 4.0 (Artificial Analysis) | 52% | 59% | Astra |
| Artificial Analysis Intelligence Index | 66 | 61 | Fable |
| Security | |||
| Prompt injection, vendor-reported | 1.0% attack success after 15 attempts | 8.5% attack success after 15 attempts | Fable |
| Acting past instructions, vendor-reported | Under 0.01% of monitored completions | 1.4% unauthorized external communication with confirmation policy | Different tests, not comparable |
| Published workplace-agent safety table | No | Yes, 8 categories, with and without confirmation policy | Astra on transparency |
| Bottom line | |||
| Recommendation | Keep for research, drafting, and review | Consider for calculation-heavy API pipelines | |
“[Practitioners] must thoroughly review all AI-created documents and language incorporated into writings before delivery to a client or submission to the IRS.”
With extensions due on September 15 and October 15, firms have a lot on their plates already. Plus, independent benchmarks on two models that remain only days old will take a while to properly calculate. The most important question for firms is who gets access to the models, but once that’s settled, choosing a model is compelling to revisit throughout the year.
Truss is the more-in-one tax workflow platform — helping accounting firms collect client info, manage workpapers, prep returns with AI support, and deliver everything in one place. Book a demo.