Forwards or Backwards

Topics › Work and the Future of Jobs

AI Agents and Humans Working Together

Published · 30 min

Watch on YouTube

An AI agent is a loop: it is given a goal and a set of tools, it takes a step, reads the result, and takes the next one - often dozens of times before a person sees anything. The popular account says that hands the work to the machine. The evidence says something less comfortable: the work does not leave. It moves to the person checking it.

This film goes through what has actually been measured.

How long a task an agent can carry: METR's 50% time horizon - the length of task, timed on human experts, that an agent clears about half the time. In METR's Time Horizon 1.1 update of 29 January 2026, across a suite of 228 tasks, Claude Opus 4.5 was measured at about five hours twenty minutes, with a confidence interval from under three hours to over twelve, and the doubling time measured from 2024 was about 89 days.

Then the warning that predates all of it: Lisanne Bainbridge's Ironies of Automation, Automatica, November 1983. Automate what you can, and what is left for the human is the rare and difficult part - plus watching. People are poor at watching. Skills fade when unused. The more advanced the control system, the more crucial the human operator.

Then the field evidence. A customer support operation, 5,179 staff: +14% issues resolved per hour on average, +34% for the least experienced, minimal for the most experienced (Brynjolfsson, Li and Raymond, QJE 2025). 758 BCG consultants across 18 tasks: +12.2% tasks, 25.1% faster, over 40% higher quality inside the frontier - and 19 percentage points worse on a task designed to sit just outside it (Dell'Acqua and others, 2023). Sixteen experienced open-source developers who forecast a 24% speedup, measured 19% slower, and still believed afterwards they had been 20% faster (METR, 10 July 2025). And the follow-up that could not be run cleanly, because 30 to 50% of developers refused to do tasks without AI (METR, 24 February 2026).

Then a shop in Anthropic's office run by an agent for a month, and what it bought, priced and hallucinated (Project Vend, 27 June 2025). What Anthropic's usage data suggests a human in the loop is worth: 50% success at about 3.5 hours of human work through the API, against an extrapolated 19 hours in conversation (Economic Index, 15 January 2026). What 20,000 knowledge workers say they do with AI output (Microsoft Work Trend Index, 5 May 2026). What EU law now requires of the humans supervising high-risk systems - Article 14, which names automation bias directly - and the Digital Omnibus that deferred those obligations to 2 December 2027.

Disclosed on screen: Anthropic is a source here and makes the Claude models, and this script was researched and drafted with Claude's help; Microsoft is a source and sells Copilot. Independent work is set beside both.

Nothing in this film claims an agent intends, wants or understands anything, and nothing here measured employment.

Educational documentary. Not financial or investment advice.

In these topics

Tags

Chapters

  1. The job that moves
  2. What an agent actually is
  3. How long a task it can hold
  4. A warning from nineteen eighty-three
  5. The call centre
  6. The jagged frontier
  7. Sixteen developers who felt faster
  8. The experiment that could no longer be run
  9. A shop run by an agent
  10. The conversation as a correction loop
  11. What people say they do with the output
  12. Writing oversight into law
  13. The projects that stop
  14. The supervisor who never did the job
  15. What this film has not said
  16. The job, restated

More from Forwards or Backwards on YouTube

Sources and credits

Primary sources

  • METR, 'Time Horizon 1.1', 29 January 2026 - the 50% time-horizon definition, the 228-task suite, only 5 of 31 long tasks with human baselines, Claude Opus 4.5 at 320 minutes CI [170, 729], and doubling times of 196 days since 2019 and 88.6 days since 2024. PRIMARY. The film says 'among the highest' rather than 'the highest' because the fetched table may be partial.
  • METR, 'Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity', 10 July 2025, arXiv 2507.09089 - 16 developers, 246 issues, forecast 24% faster, measured 19% longer (CI +2% to +39%), believed 20% faster afterwards.
  • METR, 'We are Changing our Developer Productivity Experiment Design', 24 February 2026 - 57 developers, 143 repositories, 800+ tasks; -18% (CI -38% to +9%) for returning developers, -4% (CI -15% to +9%) for new ones; 30-50% of developers withholding tasks they did not want to do without AI. PRIMARY.
  • Erik Brynjolfsson, Danielle Li and Lindsey Raymond, 'Generative AI at Work', NBER Working Paper 31161 (April 2023, rev. Nov 2023), published Quarterly Journal of Economics 140(2), 2025 - 5,179 support agents, +14% issues per hour, +34% for novices, minimal for the most experienced.
  • Dell'Acqua, McFowland III, Mollick, Lifshitz-Assaf, Kellogg, Rajendran, Krayer, Candelon and Lakhani, 'Navigating the Jagged Technological Frontier', HBS Working Paper 24-013, 15 September 2023 - 758 BCG consultants, 18 tasks, +12.2% tasks, 25.1% faster, >40% quality, +43% below-average against +17% above-average, and 19 percentage points worse outside the frontier.
  • Anthropic with Andon Labs, 'Project Vend', 27 June 2025 - Claude 3.7 Sonnet running a shop for about a month from 31 March 2025: metal cubes sold below cost, $3 drinks beside a free fridge, discounts after the fact, a non-existent payment account, and the identity episode around 1 April. DISCLOSED ON SCREEN: Anthropic makes the model and ran the experiment.
  • Anthropic, 'Economic Index report: Economic primitives', 15 January 2026 (data 13-20 November 2025) - Claude.ai 52% augmentation / 45% automation / 32% directive; API 64% directive; 50% success at about 3.5 hours through the API against an extrapolated 19 hours in conversation. ESTIMATES from classified conversations on one company's traffic; the 19-hour figure is an extrapolation, and the film says so.
  • Microsoft WorkLab, 'Work Trend Index 2026: Agents, human agency, and the opportunity for every organization', 5 May 2026 - 20,000 knowledge workers in 10 countries, fieldwork 18 February to 20 April 2026; 86% treat AI output as a starting point; 50% prioritising quality control; 15x growth in active agents; 67% of modelled impact attributed to organisational factors. A SELF-REPORTED SURVEY BY A COMPANY THAT SELLS COPILOT - disclosed on screen.
  • Gartner press release, 25 June 2025, 'Gartner Predicts Over 40% of Agentic AI Projects Will Be Canceled by End of 2027' - the prediction and its three reasons, about 130 genuine agentic vendors out of thousands, and the 2028 forecasts of 15% of daily decisions and 33% of enterprise applications.
  • Lisanne Bainbridge, 'Ironies of automation', Automatica 19(6), November 1983, 775-779 - the vigilance limit of about half an hour, the deterioration of unused skills, and the closing irony.
  • Regulation (EU) 2024/1689 (AI Act) Article 14 - effective oversight by natural persons, awareness of automation bias, correct interpretation, the power to disregard, override or reverse, and a stop button. Read at artificialintelligenceact.eu.
  • Regulation (EU) 2026/1744 (Digital Omnibus on AI) - published in the Official Journal 24 July 2026, in force 27 July 2026; Annex III high-risk obligations deferred to 2 December 2027, Annex I to 2 August 2028.

Not regulated financial advice.