Botsitting: The Unpaid Labour Behind Every AI Productivity Claim

Six thousand office workers in the United States, the United Kingdom and Australia were asked, across December 2025 and January 2026, how much time artificial intelligence had given back to them. They said eleven hours a week.
Then they were asked a second question, the one almost nobody asks. How much time do you spend feeding these systems context, supervising what comes out, debugging the mistakes and cleaning up afterwards? The answer was 6.4 hours a week.
That is the whole argument in two numbers. Roughly six of every ten hours artificial intelligence appears to save are consumed by the labour of making artificial intelligence work. The survey, published in June 2026 as the Work AI Index by Glean's Work AI Institute alongside researchers from Stanford, Notre Dame, Emory, UC Berkeley, UC Santa Barbara, UNC Charlotte and University College London, gave the residual a name that has since escaped into general use: botsitting. Rebecca Hinds, who heads the institute and co-authored the report, and her colleagues defined it coldly, as “the largely unrecognized, unbudgeted, and untracked labor of making AI usable”.
Unrecognised, unbudgeted, untracked. Three adjectives doing an enormous amount of work.
On 18 August 2026, the Indian human resources publication HRKatha ran an analysis of the same phenomenon under the more homely label of bot sitting, defining it as the work employees do prompting systems, checking answers, correcting errors, supplying missing context, trying again when outputs go wrong, and deciding whether the eventual result can be trusted at all. Its sharpest line concerns measurement. The real test of efficiency, the piece argued, is not how quickly a machine produces an answer but how much human work remains before anyone is willing to trust it.
Nobody is measuring that. And what nobody measures, nobody pays for.
The Arithmetic That Stops Working at Five Minutes
Take the scenario in its simplest form. A task used to take twenty minutes of a person's own work. Now it takes five minutes of generation followed by fifteen minutes of checking, correcting and verifying. On the stopwatch, nothing has changed. On the productivity dashboard, everything has. The dashboard records five minutes of task completion, because that is the interval in which the tool was invoked and the output produced. The fifteen minutes afterwards are logged as the worker doing their job, which is what they were doing before the tool arrived.
The tool has not saved fifteen minutes. It has reclassified them.
This is a straightforward consequence of how enterprise software reports on itself. Adoption metrics count seats, prompts and sessions. They do not count the second and third attempts, the cross-check against a source document, the quiet decision not to send the thing at all. More prompts is not a measure of value. It may be a measure of the opposite: a worker who prompts a system eleven times before getting a usable answer generates eleven data points that a usage dashboard reads as enthusiasm.
The Work AI Index found a second figure that should worry anyone relying on those dashboards. Only thirteen per cent of organisations surveyed said artificial intelligence had significantly improved their performance, against eighty-seven per cent of workers who said they were using it. That gap between individual time saved and organisational performance gained is where the botsitting hours have gone. They have not disappeared. They have been absorbed into a category of labour the accounting system cannot see.
The survey also documented what happens when supervision becomes unaffordable. Sixty-nine per cent admitted to what the report calls botshitting: shipping output they had not verified, did not fully understand, or could not confidently stand behind. Forty-one per cent had sent work they would be unable to explain if questioned. Twenty-eight per cent admitted blaming the machine for their own mistakes. Workers who reported botshitting were 3.8 times more likely to be looking for another job.
That last figure is the tell. This is not laziness. It is triage under an unfunded mandate.
Lisanne Bainbridge Wrote the Manual for This in 1983
None of this is new. The person who explained it most economically did so forty-three years ago, in a five-page paper about process control in factories and power stations.
Lisanne Bainbridge published “Ironies of Automation” in the journal Automatica in 1983. Her central observation was that the designer of an automated system regards the human operator as unreliable and inefficient, and so tries to design them out, but cannot automate everything. What remains for the human is precisely the residue that could not be specified: the awkward, ambiguous, judgement-heavy fragments that defeated the engineering. The operator is left with the hardest parts of the job, stripped of the easier parts that used to keep their skills sharp, and asked to intervene in exactly the situations for which they are now least prepared.
Bainbridge was blunt about monitoring in particular. “The human monitor has been given an impossible task,” she wrote, noting that where a computer is making decisions faster and on more dimensions than a person can follow, “there is therefore no way in which the human operator can check in real-time that the computer is following its rules correctly.” She caught the training paradox too: it is ironic, she observed, to train operators in following instructions and then place them in the system to provide intelligence.
Substitute a large language model for a distributed control system and the paper reads like a memo from last week. The knowledge worker of 2026 has been handed the residue. Drafting a first version of a summary, a function, a customer reply: that was the tractable part, and it has been automated. What is left is knowing whether the draft is right, which requires knowing the domain, which was previously maintained by doing the tractable part.
The literature that followed quantified the failure modes she predicted. Raja Parasuraman and Dietrich Manzey's 2010 review in Human Factors, synthesising decades of empirical work on automation complacency and automation bias, reached conclusions that should be printed on the login screen of every enterprise assistant. Complacency emerges specifically under multiple-task load, when manual tasks compete with the automated task for attention. It appears in expert users as readily as in novices, and cannot be trained away with simple practice. Automation bias produces both errors of omission, where the person fails to act because the system did not flag a problem, and errors of commission, where the person acts wrongly because the system told them to. Neither is reliably prevented by instructions telling people to be careful.
Meanwhile the monitoring itself is not free. Joel Warm, Parasuraman and Gerald Matthews titled their 2008 Human Factors paper on the subject with unusual directness: vigilance requires hard mental work and is stressful. Sustained attention to a mostly reliable process is not a restful state between bouts of real work. It is a demanding task with measurable workload and stress costs, and performance on it degrades over time.
So the corporate framing, in which the human is elevated from doing to overseeing, describes a promotion. The human factors literature describes a transfer into a job that is cognitively expensive, psychologically taxing, unavoidably error-prone and, in most workplaces, entirely uncompensated.
Sixteen Developers Who Were Certain They Had Gone Faster
The most instructive evidence in this debate is a small randomised controlled trial with an awkward result and an unusually honest set of authors.
In 2025, the research organisation METR recruited sixteen experienced open-source developers working on repositories they personally maintained, projects averaging over 22,000 GitHub stars. It randomised 246 real issues from those repositories into two conditions: artificial intelligence tools permitted, or not permitted. Tasks averaged around two hours. Beforehand, the developers forecast that the tools would speed them up by twenty-four per cent.
They were nineteen per cent slower with the tools.
What happened next matters more than the slowdown. After completing the work, having lived through the actual elapsed time, the same developers estimated that artificial intelligence had made them twenty per cent faster. They were wrong by roughly forty percentage points about their own labour, in the direction of the tool, on tasks they had personally performed within the previous few hours.
That is the epistemological problem at the heart of every self-reported productivity statistic in this field, including the eleven hours in the Work AI Index. Generation feels fast because it is fast and visible. Verification feels like ordinary work because it is ordinary work, diffuse, arriving in fragments scattered through the day. People are demonstrably bad at summing the second category and comparing it to the first.
METR deserves credit for what it did afterwards. In February 2026 it announced it was redesigning the experiment. Its follow-up cohorts produced point estimates of a negative eighteen per cent speedup for the original developers, with a confidence interval from negative thirty-eight to positive nine per cent, and negative four per cent for newly recruited developers. The slowdown persisted in the point estimates but the intervals now straddled zero. More importantly, METR reported that developers were increasingly refusing to take part in conditions barring them from using the tools, that the pay rate had fallen from 150 dollars an hour to fifty, and that measuring elapsed time had become genuinely hard when participants ran multiple agents concurrently. “Due to the severity of these selection effects, we are working on changes to the design of our study,” the organisation wrote, adding that the true speedup among excluded developers could be considerably higher.
That is what intellectual honesty looks like, and it cuts both ways. Anyone citing the nineteen per cent slowdown as a settled fact about artificial intelligence in 2026 is overreaching. But the perception gap is the more durable finding, and nothing in the update disturbs it. The difficulty in the follow-up was not that developers had become good at estimating their own throughput. It was that the experiment could no longer isolate the variable, because people who use these tools will no longer agree to stop.
Why Checking Is Harder Than Doing
There is a structural reason verification consumes more time than intuition suggests, and it concerns the shape of machine error.
A junior colleague who does not know something produces work that signals its own uncertainty. The prose is hedged, the gaps are obvious, the citations are missing. A language model that does not know something produces work that is fluent, confident, correctly formatted and internally consistent. The error sits in a sentence that looks exactly like every true sentence around it. The cost of finding it is therefore not the cost of scanning for anomalies. It is the cost of independently establishing the truth of each load-bearing claim, which in the limit is the cost of having done the work yourself.
This is why the five-minutes-plus-fifteen arithmetic is not a transitional inconvenience that better models will erase. As accuracy rises, the frequency of error falls but the difficulty of detection rises, because a rarer error in more plausible packaging demands more sustained vigilance to catch. Parasuraman and Manzey's finding is precisely that reliable automation breeds the attentional withdrawal making occasional failure catastrophic. Better models make botsitting less frequent and more consequential at once.
The downstream costs have now been measured in money. In September 2025, BetterUp Labs and the Stanford Social Media Lab surveyed 1,004 full-time American desk workers about what they termed workslop: output with the appearance of good work but lacking the substance to advance the task. Forty per cent had received it in the preceding month, and respondents estimated that 15.4 per cent of the work reaching them fell into the category. Each incident took an average of one hour and fifty-one minutes to sort out, around twenty minutes longer than doing the work properly in the first place would have taken. Converted using respondents' own salaries, that came to roughly 186 dollars per employee per month, or more than nine million dollars a year for an organisation of ten thousand people. Managers were markedly more exposed than individual contributors, at fifty-four per cent against 38.5 per cent.
Note what workslop is in labour terms. It is botsitting skipped upstream, landing unpriced on somebody downstream. The person who declined to verify saved fifteen minutes. The person who received the output spent one hour and fifty-one. That is not productivity. It is a transfer of unpaid supervisory work between colleagues, with interest.
What Gets Measured Was Never Designed to See This
Organisations do have instruments for measuring cognitive workload. They simply do not use them for this.
The NASA Task Load Index, developed by Sandra Hart and Lowell Staveland in 1988 and still the most widely used subjective workload instrument in ergonomics, decomposes workload into six components: mental demand, physical demand, temporal demand, performance, effort and frustration. In its original form it asks respondents to weight those components against one another across all fifteen possible pairs. Hart's twenty-year retrospective in 2006 catalogued its spread across aviation, medicine and interface design.
Look at that list and notice how badly conventional workload accounting maps onto botsitting. Corporate measurement, where it exists at all, tracks temporal demand: hours worked, tickets closed, tasks completed. Mental demand, effort and frustration are exactly the dimensions supervision loads most heavily and timesheets record least. A worker who closes the same number of tickets while spending two extra hours a day in a state of alert scepticism about machine output registers, on every dashboard the employer owns, as having had an identical day.
This is where the sociology of work becomes more useful than the economics. Susan Leigh Star and Anselm Strauss published “Layers of Silence, Arenas of Voice” in the journal Computer Supported Cooperative Work in 1999, a foundational paper on invisible work in computer-supported systems. Their opening move was to insist that what counts as work is a matter of definition, not a matter of fact, and that designers routinely automate the visible portion of a process while remaining oblivious to the articulation work holding the arrangement together: the coordinating, contextualising, patching and repairing that never appears in a process diagram and therefore never appears in a budget.
Botsitting is articulation work. Supplying missing context to a model, deciding which of three plausible outputs matches what the client meant, knowing the tool has silently used a superseded policy document: this is the connective labour Star and Strauss described, now performed on behalf of a machine rather than a colleague. Their warning was that automating around invisible work does not eliminate it. It removes the vocabulary for discussing it.
The lifecycle synthesis of human-AI collaboration risks published on arXiv in August 2026 by Md Foysal Ahmed, Isaac Kobby Anni and Md Main Uddin Rony makes a compatible point in the language of risk taxonomy. Surveying evidence across healthcare, journalism, education, research, organisational decision-making and defence, the authors identify six recurring risk clusters that cut across domains: trust miscalibration, cognitive burden, accountability gap, capability erosion, goal misalignment, and AI anxiety and technostress. These are not separable engineering defects to be fixed one at a time, they argue, but interlocking sociotechnical dynamics cascading through four lifecycle stages, which is why piecemeal interventions frequently create unintended consequences. Cognitive burden and accountability gap sitting adjacent in the same taxonomy is no coincidence. They are one problem seen from the worker's side and the organisation's side.
The History of Labour-Saving Devices Is a History of Redistribution
There is a precedent for all of this, and it is not from computing.
Ruth Schwartz Cowan's “More Work for Mother”, published in 1983 and awarded the Society for the History of Technology's Dexter Prize the following year, asked why American housewives were working longer hours in 1970 than in 1870 despite a century of mechanisation. Washing machines, vacuum cleaners, gas ovens, commercial flour and refrigeration had all arrived. The hours had not fallen.
Cowan's answer was that the appliances did not eliminate labour. They redistributed it and then raised the standard it was held to. Work previously done by servants, husbands, children and commercial services was pulled back onto one person, while expectations for cleanliness, nutritional variety and laundry frequency rose to consume whatever slack the machines created. The technology was genuinely labour-saving per unit of output. Total labour rose anyway, because output expectations rose faster and the residual tasks landed on a single unpaid worker.
Read the Work AI Index against that and the shape is familiar. Eleven hours of unit-level saving, 6.4 hours of new residual labour, thirteen per cent of organisations reporting real improvement, and a rising expectation of volume absorbing the difference.
The digital version of this dynamic already has a well-documented lower tier. Mary L. Gray and Siddharth Suri's “Ghost Work”, published in 2019, described the invisible global workforce that fills the gaps automated systems cannot close: content flagging, transcription checking, data labelling, the human intervention that makes a service look seamless. Their formulation of the paradox is that the drive to eliminate human labour reliably generates new human tasks, and that those tasks are systematically hidden, poorly paid and structurally insecure. The most cited illustration remains Billy Perrigo's January 2023 investigation for TIME, which documented that OpenAI had used workers in Kenya, employed through the outsourcing firm Sama, to label descriptions of child sexual abuse, torture, self-harm and bestiality so that ChatGPT could learn to filter such material. They were paid between 1.32 and two dollars an hour.
The point for the office worker in 2026 is not that their situation is equivalent. It plainly is not. The point is that the industry has an established pattern of relying on human labour it declines to name, and that the pattern has migrated from the outsourced periphery to the salaried core. The mechanism is identical: apparent autonomy produced by human effort the accounting treats as external to the system.
Article 14 Turns You Into the Accountability Sink
Regulation has made the informal expectation of supervision into a legal duty, and in doing so has clarified who carries the risk.
Article 14 of the European Union's Artificial Intelligence Act requires that high-risk systems be designed so they “can be effectively overseen by natural persons during the period in which they are in use”. The overseeing person must be enabled to “properly understand the relevant capacities and limitations of the high-risk AI system and be able to duly monitor its operation”; to “remain aware of the possible tendency of automatically relying or over-relying on the output produced by a high-risk AI system (automation bias)”; to decide “not to use the high-risk AI system or to otherwise disregard, override or reverse” its output; and to interrupt the system through a stop button.
Read those clauses as a job description rather than a compliance obligation. They specify a role requiring domain expertise sufficient to override a machine, metacognitive awareness of one's own susceptibility to automation bias, and sustained vigilance across the operational life of the system. There is no corresponding requirement anywhere in the Act that this role be staffed, budgeted, timetabled or paid.
The timing is instructive. Those Article 14 obligations for standalone high-risk systems were originally due on 2 August 2026. The Digital Omnibus deferring them, agreed on 6 May 2026 and confirmed by member state representatives a week later, was published in the Official Journal on 24 July 2026 and entered into force on 27 July, six days before the deadline it displaced. The revised dates are settled law: 2 December 2027 for standalone high-risk systems, and 2 August 2028 for those embedded in products already covered by European Union product safety law. The Article 50 transparency duties stayed on the original schedule and took effect earlier this month.
So the expectation that a human will supervise the machine is already operating inside every workplace that has deployed one, and the legal duty to build machines that can actually be supervised has been postponed by sixteen months. That ordering places the burden on the human before it places the corresponding design obligation on the system.
Ben Green, in a 2022 paper in Computer Law and Security Review, surveyed forty-one policies mandating human oversight of government algorithms and found two connected flaws. The evidence suggests people cannot reliably perform the oversight functions the policies assume, and as a result the policies legitimise deployment of faulty systems without addressing what is wrong with them. His remedy was to shift accountability from individual oversight to institutional oversight.
Madeleine Clare Elish gave the failure mode its name. In “Moral Crumple Zones”, published in Engaging Science, Technology, and Society in 2019, she analysed accidents involving complex automated systems and observed that responsibility is routinely misattributed to the human closest to the failure, even where that human had minimal control over the system's behaviour. The crumple zone in a car absorbs impact to protect the occupant. The moral crumple zone absorbs blame to protect the integrity of the technological system, at the expense of the nearest operator.
Twenty-eight per cent of Work AI Index respondents admitted blaming artificial intelligence for their own mistakes. Elish's argument is that the institutional traffic runs overwhelmingly the other way.
The Skills You Stop Using Are the Skills You Need to Check With
Bainbridge's most uncomfortable prediction was that operators would lose the competence they needed for the interventions automation reserved for them. Medicine has now produced the cleanest demonstration.
In October 2025, The Lancet Gastroenterology and Hepatology published a multicentre observational study of endoscopist deskilling drawn from four Polish centres in the ACCEPT trial, which introduced computer-aided polyp detection at the end of 2021 and then randomised subsequent colonoscopies to proceed with or without artificial intelligence assistance. Nineteen experienced endoscopists, each with more than two thousand colonoscopies behind them, were studied. Their adenoma detection rate in unassisted colonoscopies fell from 28.4 per cent before exposure to the tool to 22.4 per cent afterwards, an absolute decline of six percentage points.
These were not trainees. They were highly experienced clinicians whose unaided performance degraded measurably after routine assistance, on a metric directly linked to cancer prevention.
The implication for botsitting is recursive and unpleasant. The supervisory role exists because the human is supposed to catch what the machine gets wrong. That requires the domain judgement previously maintained by performing the task unaided, which is exactly what the tool has removed. The capability erosion cluster in the arXiv risk taxonomy captures the structure: the intervention creating the need for oversight simultaneously degrades the capacity to provide it. Any organisation counting on human verification as its safety net is depending on a resource its own deployment strategy is quietly consuming.
Where the Productivity Actually Landed
The macro evidence is not that artificial intelligence does nothing. It is that the gains are real, narrow, and much smaller in aggregate than the discourse implies.
The strongest firm-level result remains the study by Erik Brynjolfsson, Danielle Li and Lindsey Raymond in the Quarterly Journal of Economics in 2025, examining the staggered rollout of a generative conversational assistant across 5,172 customer support agents. Access to the assistant raised issues resolved per hour by fifteen per cent on average. Crucially, the effect was concentrated among novice and lower-skilled workers, with minimal or slightly negative effects for the most experienced. The tool distributed the accumulated tacit knowledge of the best performers to everyone else.
That result is genuine and it is also specific. Customer support has high task volume, structured interactions, immediate feedback and low per-instance verification cost. It is the best case, and not obviously the case for legal drafting, engineering, medical documentation or financial analysis, where verifying an output can cost more than producing it and the consequences of an unverified error arrive months later.
The national statistics tell a thinner story. The US Bureau of Labor Statistics reported on 6 August 2026 that nonfarm business labour productivity rose at an annualised rate of 1.4 per cent in the second quarter of 2026, and 2.2 per cent measured against the second quarter of 2025, with output up 2.5 per cent and hours worked up 0.2 per cent over the year. Unit labour costs rose at an annualised 1.3 per cent in the quarter. These are respectable figures. They are not the signature of a technology that has removed eleven hours a week from the working lives of eighty-seven per cent of office workers.
The Upwork Research Institute, surveying around 2,500 people across the United States, United Kingdom, Australia and Canada in the spring of 2024, found the gap in its rawest form. Ninety-six per cent of C-suite leaders expected artificial intelligence to raise productivity. Seventy-seven per cent of employees using it said it had increased their workload. Forty-seven per cent did not know how to achieve the productivity gains their employers expected. One in three said they were likely to quit within six months because of burnout. Kelly Monahan, managing director of the institute, framed the conclusion carefully: it is possible for the technology to raise productivity and improve well-being simultaneously, but that outcome requires a fundamental change in how work and talent are organised, not merely a change in tooling.
Two years on, the Work AI Index suggests the reorganisation has not happened. What has happened is that the residual labour acquired a name.
Counting the Hour That Was Never Saved
Naming a thing is a precondition for measuring it, and measuring it is a precondition for paying for it. Here is what measurement would actually involve.
The first step is to stop treating verification as a discretionary activity performed by conscientious individuals and start treating it as a defined task with an estimated duration. That means workload models built on generation time plus verification time, with the second term populated from observation rather than optimism. The concept already exists. Work presented at the 2026 CHI conference on human factors in computing systems by Guangrui Fan, Dandan Liu, Lihu Pan and Rui Zhang constructed a behavioural verification-load index for programmers from observable signals including compile and test failures, code churn, pauses and context switches, and showed across sixty participants that it tracked both subjective burden and correctness. Notably, in that study the tools reduced measured workload and time on task while the verification-load metric still predicted accumulating stress and fatigue over repeated use. Both things can be true. The gains are real and the residual burden is real, and only one of them currently appears in any management report.
The second step is to use an instrument designed for the purpose. The NASA Task Load Index has been in continuous use for nearly four decades, is free, takes minutes to administer, and measures precisely the dimensions supervision loads and timesheets miss. There is no methodological obstacle to running it before and after deployment of a workplace assistant, only an incentive obstacle: the results might contradict the business case.
The third step is architectural. The paper on safe and responsible artificial intelligence agents submitted to arXiv in January 2026 by Edward Cheng, Jeshua Cheng and Alice Siu argues for a three-pillar model grounded in transparency, accountability and trustworthiness, and for staged autonomy achieved through progressive validation rather than immediate full automation, by explicit analogy with the incremental rollout of autonomous driving. Reviewing the state of the art, the authors point to Magentic-UI, an open-source interface platform for human-in-the-loop agentic systems, as an example of oversight embedded through structured, repeatable mechanisms: co-planning, co-tasking, action approval and answer verification. The significance for botsitting is that these are discrete, nameable, loggable events. An oversight step that exists as a defined interaction in software can be counted, timed, staffed and, if anyone chooses, paid for. Oversight existing only as a cultural expectation that someone will check cannot.
The fourth step is to treat oversight capacity as something an organisation builds rather than something it assumes. Yao Xie and Walter Cullen argued in a December 2025 arXiv paper that major ethics guidelines and laws, the EU AI Act explicitly included, call for effective human oversight without defining it as a distinct and developable capacity. Their proposal situates it within a well-being efficacy framework integrating artificial intelligence literacy, ethical discernment and awareness of human needs, on the grounds that people inevitably project desires, fears and interests onto these systems and that oversight therefore requires the competence to examine and, where necessary, restrain problematic demands.
That identifies the right object. Human oversight is not a checkbox and it is not a personality trait. It is a skill that decays without practice, costs energy to exercise, and is currently being demanded at scale from people who have received no training in it and no allowance for it.
The fifth step determines whether any of this reaches a pay packet, and it is contractual rather than technical. The mechanism by which invisible labour has historically become visible is not measurement alone. It is bargaining. Ben Green's argument for shifting from individual to institutional oversight points the same way: the question is not whether a given worker checked carefully enough, but whether the institution deploying the system created conditions under which careful checking was possible.
Which brings the argument back to the twenty minutes. If the task now takes five minutes of generation and fifteen of verification, the honest description is that the technology has changed the composition of the work without changing its duration, and has shifted its character from production, which most people find satisfying, to inspection, which the vigilance literature has shown for decades to be tiring, stressful and prone to exactly the errors it exists to prevent. That might still be worth doing. Verification scales in ways expertise does not, and the customer support evidence shows real gains where the economics line up.
But the hour supposedly saved is not available for redeployment if a substantial fraction of it was never saved. Counting it as though it were is not an optimistic forecast. It is a measurement error, and the people absorbing the difference are the ones holding the mouse at half past six, reading a paragraph for the third time, trying to work out whether the confident sentence in the middle of it is true.
References
- Liji Narayan, “Bot sitting: When humans end up babysitting machines,” HRKatha, 18 August 2026. https://www.hrkatha.com/features/hr-pops-features/bot-sitting-when-humans-end-up-babysitting-machines/
- Work AI Institute, Glean, “The Work AI Index 2026,” June 2026. https://www.glean.com/work-ai-institute/reports/work-ai-index
- METR, “Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity,” 10 July 2025. https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/
- METR, “We are Changing our Developer Productivity Experiment Design,” 24 February 2026. https://metr.org/blog/2026-02-24-uplift-update/
- Upwork Inc., “Upwork Study Finds Employee Workloads Rising Despite Increased C-Suite Investment in Artificial Intelligence,” 23 July 2024. https://investors.upwork.com/news-releases/news-release-details/upwork-study-finds-employee-workloads-rising-despite-increased-c
- Lisanne Bainbridge, “Ironies of Automation,” Automatica, vol. 19, no. 6, 1983, pp. 775-779. https://www.sciencedirect.com/science/article/abs/pii/0005109883900468
- Raja Parasuraman and Dietrich H. Manzey, “Complacency and Bias in Human Use of Automation: An Attentional Integration,” Human Factors, vol. 52, no. 3, June 2010, pp. 381-410. https://journals.sagepub.com/doi/10.1177/0018720810376055
- Joel S. Warm, Raja Parasuraman and Gerald Matthews, “Vigilance Requires Hard Mental Work and Is Stressful,” Human Factors, vol. 50, no. 3, June 2008, pp. 433-441. https://journals.sagepub.com/doi/10.1518/001872008X312152
- Sandra G. Hart, “NASA-Task Load Index (NASA-TLX); 20 Years Later,” Proceedings of the Human Factors and Ergonomics Society Annual Meeting, vol. 50, no. 9, October 2006, pp. 904-908. https://journals.sagepub.com/doi/10.1177/154193120605000909
- Susan Leigh Star and Anselm Strauss, “Layers of Silence, Arenas of Voice: The Ecology of Visible and Invisible Work,” Computer Supported Cooperative Work, vol. 8, nos. 1-2, March 1999, pp. 9-30. https://dl.acm.org/doi/10.1023/A:1008651105359
- Ruth Schwartz Cowan, More Work for Mother: The Ironies of Household Technology from the Open Hearth to the Microwave, Basic Books, 1983. https://www.hachettebookgroup.com/titles/ruth-schwartz-cowan/more-work-for-mother/9780465047321/
- Mary L. Gray and Siddharth Suri, Ghost Work: How to Stop Silicon Valley from Building a New Global Underclass, Houghton Mifflin Harcourt, 2019. https://ghostwork.info/
- Billy Perrigo, “Exclusive: OpenAI Used Kenyan Workers on Less Than $2 Per Hour to Make ChatGPT Less Toxic,” TIME, 18 January 2023. https://time.com/6247678/openai-chatgpt-kenya-workers/
- BetterUp Labs and Stanford Social Media Lab, “Workslop: The Hidden Cost of AI-Generated Busywork,” September 2025. https://www.betterup.com/blog/hidden-costs-workslop
- Erik Brynjolfsson, Danielle Li and Lindsey R. Raymond, “Generative AI at Work,” The Quarterly Journal of Economics, vol. 140, no. 2, May 2025, pp. 889-942. https://academic.oup.com/qje/article/140/2/889/7990658
- “Endoscopist deskilling risk after exposure to artificial intelligence in colonoscopy: a multicentre, observational study,” The Lancet Gastroenterology and Hepatology, October 2025. https://www.thelancet.com/journals/langas/article/PIIS2468-1253(25)00133-5/abstract
- European Union, Regulation (EU) 2024/1689, Article 14: Human Oversight. https://artificialintelligenceact.eu/article/14/
- U.S. Bureau of Labor Statistics, “Productivity and Costs, Second Quarter 2026, Preliminary,” 6 August 2026. https://www.bls.gov/news.release/archives/prod2_08062026.htm
- Gibson Dunn, “EU AI Act Omnibus Agreement: Postponed High-Risk Deadlines and Other Key Changes,” 2026. https://www.gibsondunn.com/eu-ai-act-omnibus-agreement-postponed-high-risk-deadlines-and-other-key-changes/
- Ben Green, “The Flaws of Policies Requiring Human Oversight of Government Algorithms,” Computer Law and Security Review, vol. 45, 2022. https://papers.ssrn.com/sol3/papers.cfm?abstract_id=3921216
- Madeleine Clare Elish, “Moral Crumple Zones: Cautionary Tales in Human-Robot Interaction,” Engaging Science, Technology, and Society, vol. 5, 2019, pp. 40-60. https://estsjournal.org/index.php/ests/article/view/260
- Md Foysal Ahmed, Isaac Kobby Anni and Md Main Uddin Rony, “Toward Resilient Human-AI Collaboration: A Lifecycle Taxonomy of Sociotechnical Risks and Cascading Failures,” arXiv:2608.05614, 6 August 2026. https://arxiv.org/abs/2608.05614
- Edward C. Cheng, Jeshua Cheng and Alice Siu, “Toward Safe and Responsible AI Agents: A Three-Pillar Model for Transparency, Accountability, and Trustworthiness,” arXiv:2601.06223, 9 January 2026. https://arxiv.org/abs/2601.06223
- Yao Xie and Walter Cullen, “Beyond Procedural Compliance: Human Oversight as a Dimension of Well-being Efficacy in AI Governance,” arXiv:2512.13768, 15 December 2025. https://arxiv.org/abs/2512.13768
- Guangrui Fan, Dandan Liu, Lihu Pan and Rui Zhang, “When Help Hurts: Verification Load and Fatigue with AI Coding Assistants,” Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems. https://doi.org/10.1145/3772318.3791176

Tim Green UK-based Systems Theorist & Independent Technology Writer
Tim explores the intersections of artificial intelligence, decentralised cognition, and posthuman ethics. His work, published at smarterarticles.co.uk, challenges dominant narratives of technological progress while proposing interdisciplinary frameworks for collective intelligence and digital stewardship.
His writing has been featured on Ground News and shared by independent researchers across both academic and technological communities.
ORCID: 0009-0002-0156-9795 Email: tim@smarterarticles.co.uk
Listen to the free weekly SmarterArticles Podcast