How we tested
We took five tasks a small-business owner really gives an AI assistant, and for each one wrote down the correct answer in advance:
- An email to a customer whose gift missed her daughter's birthday. We check the date, the tone, and whether the assistant invents facts or compensation.
- A product description — a children's slide — within 600 characters, from six specifications. We check that all six are there and nothing is made up.
- Risks in a supply contract. We planted four traps for the supplier in four clauses.
- Average order value and margin over two months. The correct figures are known.
- The 2026 income limit for a group-3 sole trader in Ukraine and what happens after exceeding it in May. The reference is our article on the sole-trader limit, checked against the Tax Code.
Three models from two companies took part: OpenAI's paid flagship GPT-5.5 and two fast models — OpenAI's GPT-5.4 mini and Google's Gemini 3.5 Flash. All answered via API, without web search, with default settings, one attempt per task.
Two important contenders are missing. We did not test Anthropic's Claude: the editorial team uses it, and scoring it against our own criteria would be a conflict of interest. Gemini 3.1 Pro is not available on the free API tier, so Google is represented by its fast model.
Results
✓ — matches the reference; ≈ — task done, but with an error you would have to fix by hand; ✗ — the answer cannot be used without rework.
All three did the calculation without errors: average order of UAH 450 and 420, margin of 40% and 39%, receipts down 12%. All three also spotted the main point — the order value grew while customers fell away. It is the same case we covered in the article on average order value.
Tax: where the cheaper models fail
The 2026 income limit for a group-3 sole trader is UAH 10,091,049: 1,167 minimum wages of UAH 8,647. GPT-5.5 named exactly that figure, with the correct consequences: 15% tax on the excess and a switch to the general system from 1 July.
The two fast models gave other figures:
- GPT-5.4 mini — UAH 8,285,700. The formula is right, but the minimum wage used is UAH 7,100, the 2024 level. It also advised "moving to another group", although there is no group above the third.
- Gemini 3.5 Flash — UAH 9,336,000, with a minimum wage of UAH 8,000, adding that the 2026 budget had not yet been adopted. The other consequences — 15%, 1 July, the application by 20 July — it described correctly.
The mechanism is simple: the model answers from memory, and its memory stops at its training date. It knows the formula but not this year's minimum wage. So the answer sounds convincing and uses the right words, but the figure is out of date. For a business owner that is the most dangerous kind of error: you cannot see it without checking.
Invented details in customer-facing texts
In the customer email, Gemini 3.5 Flash wrote that the parcel had "got stuck" and that the reshipment was "at our expense" — neither fact was in the brief. It also suggested adding a discount or a gift, though the shop had promised nothing of the kind. The suggestion is marked as optional, but a 903-character reply instead of a short one has to be rewritten before sending.
In the product description, the same model called the lacquer "hypoallergenic". The brief said only "water-based lacquer". The OpenAI models added a softer "safe for children" — also a judgement not in the data, but not a specific property someone could test.
For a business the difference matters. A promise in a customer email becomes an obligation, and a property in a product description becomes grounds for a complaint.
Contracts: the traps are found, the laws are cited wrongly
We planted four traps for the supplier: payment in 90 banking days, a penalty of 1% a day on the value of the whole contract with no ceiling, disputes heard in the buyer's court, and automatic renewal with 90 days' notice. All three models found all four and proposed sensible edits.
But Gemini 3.5 Flash added two legal claims that do not survive checking:
- it cited the Commercial Code of Ukraine, which ceased to be in force on 28 August 2025 under Law No. 4196-IX;
- it wrote that the penalty for late delivery is capped at double the National Bank's discount rate. That cap comes from the law on liability for late performance of monetary obligations, so it applies to late payment, not late delivery.
The OpenAI models did not cite specific articles of law. GPT-5.5's answer was the most complete, but also the longest — almost 8,000 characters for four clauses.
How to use this
- Calculations and structured tasks can go to any of the three models. Spot-check them.
- Figures that change every year — limits, rates, the minimum wage — never take from a model's answer without the primary source. If your assistant has web search, turn it on for such questions, and even then check against the tax authority's site or the text of the law.
- Customer-facing texts — reread with one question: has a promise or a fact appeared that you did not give?
- Contracts — use the assistant as a risk checklist, not as a lawyer. Check references to laws on zakon.rada.gov.ua, which shows whether a rule is in force.
Our position
For a small business the difference between a paid flagship model and a free fast one is not in the writing style but in whether you can trust current facts. If you ask an assistant about tax, law or government programmes, a paid model with web search pays for itself the first time it saves you from a mistake in a tax return.
The position does not hold if you use the assistant only for drafts, calculations and correspondence where you supply all the facts yourself. There, in our test, the fast models do just as well: GPT-5.4 mini handled all five tasks in 23 seconds, GPT-5.5 in 83.
Method and prompts
The test was run on 6 October 2026. Models: gpt-5.5-2026-04-23, gpt-5.4-mini-2026-03-17 (OpenAI Responses API) and gemini-3.5-flash (Gemini API). Gemini 3.8 Flash did not respond on the day due to overload; Gemini 3.1 Pro is not available on the free API tier. One attempt per task, no system instructions, no web search. Answers in the ChatGPT and Gemini apps may differ: they have search and their own settings.
Scoring rests on verifiable criteria: correct figures, the number of planted risks found, and facts that were not in the brief. The status of laws was checked on zakon.rada.gov.ua the same day.
The prompts were in Ukrainian; their full text is given in the Ukrainian version of this article.
