Ask some of us here what worries us about AI and investing and it isn't bubbles, job losses or machines gone rogue. AI will probably make us dramatically more productive at almost everything we do in research, and yet, our investment decisions may not get any faster. In fact, without redesign, they could get worse. A parallel from 60-year-old economics gives this pairing a name.
The babysitter problem
In the mid-1960s, William Baumol and William Bowen set out to explain why America's orchestras and theatres were always broke. Their answer started with the string quartet. A performance requires the same four musicians and the same half hour as in Beethoven's day. Manufacturing gets relentlessly more productive. Live music does not, yet musicians' wages rise anyway, set in an economy-wide labour market. So anything that doesn't get more productive gets relatively more expensive, forever. Baumol called it the cost disease. It is why your laptop costs less every year and your lawyer costs more.
Nordhaus put the limit case memorably:
“For increasing capabilities of computers to lead to the Singularity would require that AI could encompass all human activities … but also lay hands on patients, babysit and comfort children, and mediate disputes.”
— William Nordhaus (2021)
Figure 1 plots 125 US industries over nearly four decades of government data. The more productive an industry became, the slower its prices rose (rank correlation −0.50). Some of that is noise, and some is accounting, since prices are largely passed-through costs. The extremes tell the story. Over 1987–2023, computer manufacturing improved productivity by around 11% a year while its prices fell 10% a year. Legal services, judgment-bound work we will meet again, recorded no productivity growth, with prices up about 4% a year. As Aghion, Jones and Jones put it, growth may be constrained “not by what we do well but rather by what is essential and yet hard to improve”.
Figure 1. US growth in output prices versus growth in total factor productivity by industry, 1987–2023
Source: Man Group calculations and Bureau of Labor Statistics, Office of Productivity and Technology: major industries, 1987–2024 (released 19 March 2026); detailed industries in manufacturing, air transportation and line-haul railroads, 1987–2023 (released 26 August 2026); combined panel through 2023.
Problems loading this infographic? - Please click here
The same shape, arriving in research
Generating hypotheses, writing code and running back-tests are the computer-manufacturing side of the business, and AI is collapsing their cost as we speak. Candidate ideas are becoming abundant.
But between a candidate idea and client capital sits the judgment-bound work of deciding what to trust, the babysitter problem in miniature. That decision consumes fresh market history and calendar time. Markets produce one year of evidence per year no matter how many strategies you test, and the more candidates you test, the higher the bar of proof each must clear.1 Evidence is the string quartet; validation is the legal services of the research process.
We believe AI should cheapen parts of validation too, like better experiment design and sharper triage. But the binding input is not computation. An intraday strategy mints thousands of fresh observations a year. A macro strategy with a six-month holding period has produced a few dozen in modern market history. And the slow end is where size has to live (fast strategies hit capacity limits), so the horizons with the least evidence carry the most capital. Breadth adds correlated copies of the same few regimes (2008 arrived everywhere at once), and a simulator can stress a strategy without certifying one. Compute cannot manufacture history; only the calendar can.
Baumol's version ran on wages, whereas the research version runs on evidence, a fixed stock against a growing queue of claims. Either way, the stage that cannot improve absorbs the budget. Without redesign, validation becomes the rate-limiting and most expensive step, or it gets done badly and decisions get worse even as everything else gets faster. An early sign of the “worse” is on record, from business operations rather than investing. In Harvard Business Review, Ferreira and Tong report people applying one flat level of trust to AI (overriding when they shouldn't, deferring when they should) and doing worse in both directions. Figure 2 sketches where that can land.
Figure 2. AI accelerates ideation and implementation; without redesign, validation risks becoming the rate-limiting and most costly step, or it gets done badly, and decisions get worse

Source: Man Group. Schematic for illustrative purposes; not measured data.
So what happens next?
The excitement gathers around idea generation. The constraint, though, is moving to evidence.
We believe the response is to budget evidence the way the industry learned to budget risk. Inside our systematic business, automated experiments declare success bars and kill conditions before any result is seen, and research agents promote ideas within explicit evidence budgets. We are early. There have been two runs, both in the same market. In the first, an AI agent proposed 260 candidate signals; 20 cleared a bar set before the tests ran, on held-out history, and now trade in production,2 a 7.7% survival rate. The second proposed 338, of which 22 survived, or 6.5%. That holdout is a triage step, not the proof itself. Testing on historical data is the tail end of ideation. Those survivors join the queue for live market time, the evidence that is actually scarce.
Every firm an allocator funds faces some version of the same arithmetic. The question for a manager is shifting from how many ideas its AI can produce to how it rations evidence. That means asking how many independent out-of-sample periods stand behind a strategy, whether tests were registered before results were seen and how live returns compare with the simulation that sold the idea. Decisions tend to get worse where that discipline is missing. Nordhaus asked what a babysitter will cost when computers can do everything else. A babysitter is who you trust with what you cannot afford to lose. Inside an investment firm, that is the work of deciding which ideas deserve client capital. It runs on evidence and calendar time, and that likely means its price is going up.
Authors: Gregory Bond, Chief Investment Officer at Man Group and Gary Collier, Chief Technology Officer at Man Group.
1. A winner selected from a larger search is more likely to have benefited from noise, so the bar has to rise with the size of the search. Statisticians have formal corrections for this (Bonferroni and its successors); what matters is the number of genuinely independent candidates, and with enough of them, a back-test edge of ordinary economic size becomes indistinguishable from luck. That raises the value of ex-ante economic priors, which are a product of judgment before any test is run, not of the process that mass-produces candidates. Alongside proprietary evidence and research governance, those priors become a differentiator between firms.
2. Internal figures, September 2026; process statistics, not performance results. Production status is a deployment fact and implies nothing yet about live returns. First run: 260 candidates, 25 passed in-sample, 20 out-of-sample. Second run: 338, 56, 22. The second run's in-sample screen let more candidates through, and the held-out test then removed far more of them (20 of 25 survived the holdout in the first run; 22 of 56 in the second), consistent with the agent being pushed toward hypotheses beyond those already explored, and a reminder that the test on unseen data is the screen doing the real work.
Further reading
“The Productivity Paradox: When Will AI Deliver?” (Man Insights, February 2026); “AI Boom, Bust or Something Else? Citrini's Future” (with Panashe Bera, Oxford Man Institute, March 2026); “Will AI Make Firms Bigger or Smaller?” (Man Insights, August 2026).
References
Baumol, W. and W. Bowen (1966), Performing Arts: The Economic Dilemma, Twentieth Century Fund
Baumol, W. (1967), “Macroeconomics of Unbalanced Growth: The Anatomy of Urban Crisis”, American Economic Review 57(3)
Nordhaus, W. (2021), “Are We Approaching an Economic Singularity? Information Technology and the Future of Economic Growth”, AEJ: Macroeconomics 13(1)
Aghion, P., B. Jones and C. Jones (2019), “Artificial Intelligence and Economic Growth”, in The Economics of Artificial Intelligence, University of Chicago Press
Ferreira, K. J. and J. Tong (2026), “How AI Agents Orchestrate Work Across Silos”, Harvard Business Review, September–October 2026
You are now leaving Man Group’s website
You are leaving Man Group’s website and entering a third-party website that is not controlled, maintained, or monitored by Man Group. Man Group is not responsible for the content or availability of the third-party website. By leaving Man Group’s website, you will be subject to the third-party website’s terms, policies and/or notices, including those related to privacy and security, as applicable.