For the person who has to justify a decision about AI to a class, a team, a board or a parent. The four questions you will be asked, with sourced answers and dates you can name out loud.
- The best agents complete roughly one professional task in four at first attempt, and around 40% after eight attempts.
- Agent success on a computer-use benchmark went from 12% to 66% in a year — then the benchmark saturated, a harder one replaced it, and the best model scored 20.6%.
Everything written about AI is written for the person using it. This is for the person who has to explain a decision about it to someone else — a class, a team, a board, a parent, a client.
The four questions you will actually be asked
"Is it going to take my job?"
Answer with the measurement, not the forecast. The best agents complete roughly one professional task in four at first attempt, and around 40% after eight attempts.Mercor APEX-Agents benchmark, 21 Jan 2026 Autonomous deployment across business functions remains in single digits, and Gartner expects more than 40% of agentic projects to be cancelled by the end of 2027.Stanford HAI 2026 AI Index · Gartner, 2026
What that supports saying: it changes tasks faster than it changes jobs, and the gap between demonstration and deployment is currently very wide.
"Can we just detect it?"
No, and saying so protects you. Seven detectors tested on 91 essays falsely flagged 61.2% of work by non-native English writers, against near-perfect accuracy on native writers.Liang, Yuksekgonul, Mao, Wu & Zou, Patterns 4(7):100779, July 2023
What that supports saying: a detector output is not evidence, and using one to accuse a person is a decision you would have to defend. The full case is on how to spot AI writing.
"Why did it make that up?"
Because it was scored for guessing. Benchmarks award full credit for a lucky guess and zero for "I do not know," so a system optimised against them learns to answer everything.Kalai, Nachum, Vempala & Zhang, OpenAI, September 2025 (arXiv 2509.04664)
What that supports saying: it is not malfunctioning, it is doing what it was rewarded for. Full mechanism on why AI makes things up.
"Are we allowed to use it?"
In the EU and UK, since 2 August 2026, transparency obligations are enforceable. Article 50(4) requires labelling of AI-generated text published to inform the public only where no human editorial control was exercised — so work a person commissioned, checked and owns falls outside it.Regulation (EU) 2024/1689, Art. 50(4) · Commission Guidelines, 20 Jul 2026
What that supports saying: the obligation attaches to unreviewed output, not to assistance. More on the rules arriving.
Three things that hold a room
- Name the source out loud. "Stanford, 2023, seven detectors, ninety-one essays" ends an argument that "studies show" does not.
- Give the date. Half the disagreements in these conversations are two people holding figures from different years.
- Say what you do not know. The person who admits the uncertain part is trusted on the certain part.
Agent success on a computer-use benchmark went from 12% to 66% in a year — then the benchmark saturated, a harder one replaced it, and the best model scored 20.6%.
Nothing about the models changed. The instrument did. That single fact explains most of the disagreement people have about AI progress, and it is on the agents page with its sources.
What not to say
- Not "90% of the internet will be AI by 2026." A 2022 Europol projection, repeated for years, that did not happen.Europol Innovation Lab, Facing Reality?, 2022 — a projection, not a measurement
- Not "AI detectors are 99% accurate." Vendor marketing; independent testing does not support it.
- Not a single figure for AI energy use. Credible estimates differ by an order of magnitude; a confident number is a sign the speaker has not read the range.
Every site in this category teaches you what to do. When not to use AI covers the other side — where these tools are worse than doing it yourself, and how to tell in advance.
It is the most differentiated page on this site, and it is the one people send to other people.