TL;DR: Yesterday I put this question to the CFA Institute Sustainable Investing Community alongside my co-authors Philip Sun and Jared Broad. My answer: yes, AI trading has an ESG problem — three of them — unless you stop treating ESG as a label to buy and start using it as three lenses to look through. Ignore the Environmental lens and a key input reprices under you: conventional DRAM contract prices rose roughly 90–95% in one quarter this year. Ignore the Social lens and you build key-person risk in reverse: the agent skips the check the junior never learns. Ignore the Governance lens and models overrun you: a thousand backtests a night with nobody who can stop the agent at 2 a.m. The fix is a ledger with three blanks — Receipt, Person, Stop — that a manager either can or cannot fill. The full deck is embedded below.

A RAM stick, a receipt with times and amounts under a magnifier, and a large red button on a blue surface.

First, the thanks. George Marrash and Richard McGillivray of CFA Institute organised yesterday’s session as an event of the Sustainable Investing Community, and it was their idea to put a quant, a hedge-fund veteran and a platform founder in front of an ESG audience rather than in a technology track. Philip Sun walked through what AI actually does at each stage of a trading loop; Jared Broad ran a live multi-agent research demo from the book’s code with compute and cost on screen; I did the framing and the ledger. The recording is available to community members on the event page. If you missed the pre-event outline, this post is the payoff.

Does AI trading have an ESG problem?

Yes. And not the one the question implies.

The question assumes “ESG” means a moral verdict — is this desk good or bad for the planet, for people, for the rules? On that reading, the answer is a shrug: a trading desk’s kilowatt-hours are cents against the financed emissions of its portfolio, and nobody is going to prison for a backtest.

Read ESG the way a risk manager reads it — as three lenses, each pointed at a class of risk you are already carrying — and the answer flips. AI trading has an E problem (a production input with no receipt), an S problem (nobody left who is competent when the agent is wrong) and a G problem (no named person who can stop it, running on the same two providers as everyone else). Every one of them is P&L. None of them is measured at the unit that matters.

That is the sentence I opened with and the one I will keep repeating until it stops being controversial: sustainability is not a conscience. It is P&L you have not measured yet.

Mental model: inversion. Do not ask “is this manager sustainable?” Ask “what would a manager who is not be unable to show me?” The list is short and specific. No receipt — watt-hours per accepted decision. No person — human minutes per accepted decision. No stop — one named owner who can halt the agent at 2 a.m. If a desk cannot fill those three cells, that is a diligence signal, not a moral one.

Why should a desk that doesn’t believe in ESG use the lenses anyway?

Because the people who sell you compute, allocate you capital, clear your trades and license your fund already look through them, and each of them is a line on your own P&L. In the deck I called these the four reasons to care — cost, capacity, capital and counterparties, contagion — and noted that none of them is moral.

The mistake most desks make is to treat this as a cost-only question. Cost is where a desk starts. The forward curve is where it ends. I showed three bands:

BandWhat is in itExample
Priced todayCompute, memory, storage, colocation, power tariff and the carbon tax already inside the tariff. FinOps territory.Singapore’s carbon tax is S$45/tCO2e for 2026–27 on a stated path to S$50–80 by 2030 (NCCS). At Singapore’s grid factor that is roughly S$0.018 of carbon in every kWh today — my derivation, not a regulator’s print.
Being repriced 2026–2030Hourly Scope 2, provider-level model energy documentation, climate plans in your own accounts. Proposals and guidance today; line items inside one fund cycle.GHG Protocol hourly Scope 2 consultation; EU AI Act Annex XI documentation duties on model providers; EU data centres ≥500 kW reporting energy, water and PUE.
Not on an invoiceGrid queue, water in stressed basins, embodied carbon, model concentration, systemic correlation. Shows up as scarcity, refusal and loss, not as line items.Two providers carry roughly two-thirds of enterprise LLM API spend (Menlo Ventures, December 2025).

The arrow points up. Manage only the top band and you are hedging the 2024 version of your cost base.

A line graph with shaded areas showing future costs: Priced Today in yellow, Being Repriced 2026-2030 in orange, and Not on an Invoice in grey, with related cost factors labeled.

What happens when you ignore the E lens? Your inputs reprice without you

The Environmental lens, for a desk, is not a carbon report. It is the question: which physical inputs does my research loop consume, at what unit, and who else is bidding for them? Watts, gigabytes of RAM, GPU-hours, rack space, and the grid behind them.

Here is what 2026 did to a desk that never asked:

  • Memory. TrendForce raised its forecast for conventional DRAM contract prices to +90–95% quarter-on-quarter for Q1 2026 (TrendForce, 2 February 2026; Reuters), and its June survey put the realised increase at roughly 93–98% (TrendForce, 1 June 2026). A research loop that holds ten years of tick data in RAM because “memory is cheap” had its cost base doubled by hyperscalers it never competes with in any market except this one.
  • Accelerators. Per SemiAnalysis data I cited in the deck, one-year H100 rental went from about $1.70 to $2.35 an hour between October 2025 and March 2026 — up 40% in five months — while on-demand capacity sold out across providers.
  • Capex. US hyperscalers are forecast to spend about US$800 billion on infrastructure this year, on analyst consensus compiled by Goldman Sachs. Compute is rationed, and the biggest buyers are served first.
  • Geography. Singapore wholesale colocation runs roughly US$330–475 per kW per month against about US$110–150 in Johor (Coloprice), while Peninsular Malaysia’s grid emits about 0.740 kgCO2/kWh against Singapore’s 0.402, on the regulators’ own prints. The cheap region is the dirty region, by 1.84×. Hourly Scope 2 would make region and hour reportable.

Historical precedent: the 2011 Thailand floods. In October 2011, monsoon flooding hit factories that produced close to half the world’s hard drives; a global shortage lasted through 2012 and drive prices roughly doubled (Backblaze; Boston Globe). Every IT budget that had treated storage as a commodity discovered it had been running an unhedged short in a physical input. The 2026 DRAM shock is the same trade with a different weather system: this time the storm is US$800 billion of hyperscaler demand.

The analogy I used on the call: a desk that does not know its watt-hours per decision is a restaurant that does not know its food cost per plate. It can be profitable for years. It finds out how exposed it was the quarter that eggs double.

Mental model: Jevons paradox. Energy per AI task is falling roughly an order of magnitude a year; volume is rising faster. Cheaper compute per experiment does not mean less compute. It means more experiments. Which is why the most sustainable thing a PM can do this quarter is not buy an ESG label. It is halve the memory footprint of the research loop — which also happens to be the biggest available hedge against the DRAM shock. Cost, resilience and footprint converge on the same lever.

And the part that makes E a disclosure problem rather than a cost problem: the models on your desk will not show you the watts. Google publishes a measured 0.24 Wh median for a Gemini text prompt; Microsoft Research’s peer-reviewed study in Joule models a frontier chatbot at 0.31 Wh median rising to 3.91 Wh for long reasoning — about 13× (Microsoft Research, April 2026). The frontier models most desks actually call publish a price per million tokens and no energy figure at all. Five times the price is not five times the watt-hours — and nobody outside the provider knows what it is. “Not disclosed” is a position. It is short information.

What happens when you ignore the S lens? Key-person risk, inverted

The Social lens, for a desk, is not charity. It is bench, title and client — the people whose competence, rights and trust your process depends on — plus the market whose liquidity you consume.

A dark office with multiple rows of computer monitors displaying orange code, empty chairs, and cables on desks; city lights visible through windows in the background.

The uncomfortable arithmetic: the agent saved four days. Whose four days? And who is still competent when it is wrong?

  • Bench. The August 2026 revision of Stanford’s Canaries in the Coal Mine, using ADP payroll data through June 2026, finds employment of 22–25-year-olds in AI-exposed occupations sits 19% below less-exposed peers — descriptive, not causal (Stanford Digital Economy Lab). On a desk this shows up as key-person risk inverted: classic key-person risk is one senior who knows everything leaving. The new version is that the agent does the junior’s work, so the junior never learns the check, and in five years nobody on the floor can tell when the agent is wrong.
  • Title. Anthropic’s US$1.5 billion settlement with authors over pirated training books received preliminary court approval in September 2025 (AP News). Licensing is a contingent liability on your vendor’s balance sheet — until the feed you ingested, or the client document you pasted into an unapproved service, moves it onto yours.
  • Client. The SEC’s first AI-washing penalties — US$400,000 combined against Delphia and Global Predictions in March 2024 — were small on purpose (SEC). Mis-selling with an AI label is mis-selling. CFA Standard V(A), reasonable basis, already covers it.
  • Market. Liquidity is a social good you consume. When every agent runs the same prompt against the same model and sells, nobody is on the bid.

Historical precedent: Air France 447, June 2009. The BEA’s final report found the crew failed to identify the approach to stall after the autopilot disconnected, and cited a lack of training for high-altitude hand-flying (BEA final report; NASA safety message). The automation had been so good for so long that the manual skill it replaced had quietly atrophied. That is the S risk on an agentic desk in one sentence: the failure mode is not the agent’s error, it is the human’s inability to recognise it.

Mental model: second-order effects. The first-order effect of an agent is speed. The second-order effect is the skill that stops being practised. The desk number that captures it does not exist until you create it: human minutes per accepted decision, with an override log, and a deliberate list of apprenticeship tasks you keep doing by hand because the agent could do them.

What happens when you ignore the G lens? Models overrun you

The Governance lens, for a desk, is three questions: can I validate it, can I stop it, and what happens when everyone else is running the same thing?

  • Validation. When an agent can produce a plausible backtest in twenty seconds, the Deflated Sharpe Ratio — a Sharpe ratio corrected for the number of trials you ran, their variance, sample length, skew and kurtosis — became the only number in the industry that got more important, not less. Example: run 5,000 parameter variants on ten years of pure noise and the best in-sample Sharpe you will find is about 1.31. Report it without the trial count and you have reported nothing. An agent that runs a thousand backtests overnight does not solve the multiple-testing problem; it industrialises it.
  • Authority and stop. Who can halt the agent at 2 a.m.? Does “stop” halt new actions or flatten the book? Show me the last action your controls actually denied. The FSB’s June 2026 consultation lists 12 sound practices for responsible AI adoption — and zero statutes (FSB). More pointedly: the Federal Reserve’s SR 26-2, issued in April 2026, supersedes SR 11-7 as the US model-risk guidance and places generative and agentic AI outside its scope (Federal Reserve). The supervisory letter that used to cover your models has just told you it does not cover these. Your own policy has to.
  • Concentration. Two providers account for roughly 67% of enterprise LLM API spend (Menlo Ventures). On 20 October 2025, AWS’s us-east-1 region was disrupted for more than 15 hours (ThousandEyes). Two-thirds of the market’s model spend in two providers is not diversification. It is a single point of failure with two logos.
Diagram showing One trade wearing many logos, with icons linked to Provider A and Provider B, both pointing to a box labeled One Trade..

Historical precedent, part one: Knight Capital, 1 August 2012. A faulty deployment sent Knight’s router firing orders for 45 minutes before anyone could stop it, producing a US$440 million loss and, later, a US$12 million SEC settlement for market-access failures (SEC order; SEC press release). That was one firm’s code with humans in the building. An agentic research loop with live credentials and no rehearsed kill switch is Knight with the humans asleep.

Historical precedent, part two: the Quant Quake, August 2007. In the week of 6 August 2007, a set of quantitative long/short equity funds suffered unprecedented losses at the same time, which Khandani and Lo attributed to a coordinated unwind of crowded, similar strategies (Khandani & Lo). Nobody had a bad model. Everybody had the same model. Now imagine the same crowding with the strategies generated by the same two foundation models, in the same region, from similar prompts. CFA Institute’s own Algorithmic Market Hypothesis names this algorithmic monoculture a systemic risk (CFA Institute, 2026).

Mental model: systemic versus idiosyncratic risk. Your model being wrong is idiosyncratic; you can diversify it. Everyone’s model being the same is systemic; you cannot. The G lens is the only one of the three that asks the second question.

What does the ledger look like filled in?

Three columns labeled Receipt with a lightning bolt, Person with a stopwatch, and Stop with a pointing finger icon. Text below reads: A BLANK IS AN ANSWER..

The three lenses collapse into three blanks a manager either can or cannot fill. Binary, not rhetorical. Mental model: falsifiability.

Receipt · EPerson · SStop · G
Not small. Unmeasured at the unit that matters.Whose four days. Who is still competent when it is wrong.Your correlation risk, and your stop.
kWh per accepted decision, labelled measured / estimated / not disclosedHuman minutes per accepted decisionTrial count and the five deflated-Sharpe inputs
Campaign total beside the ratioOverride log; apprenticeship tasks retainedOne named owner; the last denied action
Region, hour and grid factorLicence status of every ingested feedKill switch vs orderly wind-down, rehearsed
RAM per run; silicon mixClient remediation route and named ownerProvider share of runs; fallback test date

Evidence rule, under all three: model + version · data known at decision time · every trial counted · point-in-time for price, ESG and the model itself. That last clause matters. Every price-data hazard — look-ahead, survivorship, restatement — applies to ESG data verbatim, and a language model that has read the 2020 restatement will rate the 2018 issuer with it. Masking the date in the prompt does not remove the weights. I have written about that failure mode as the Release-Date Rule.

Honestly labelled, what binds today: CFA Standard V(A) and the Singapore carbon tax. What is guidance: SR 26-2, with agents out of scope. What is proposal: the FSB’s 12 practices and hourly Scope 2. What is left: your own policy, covering what the letters decline.

Which six questions should you ask on your next manager call?

#QuestionNumber to ask forLabel you will acceptOwner
1What did your last accepted idea cost, including every discarded trial?$ and kWh per accepted decision; total trials; campaign totalMeasured or estimated. “Not disclosed” is an answer — record it.Operational DD
2Which silicon runs the research loop, and why?Silicon mix; ≈20% hourly saving on ARM where compatibleVendor quote is fineOperational DD
3What is RAM per run, and what breaks if memory prices double again?RAM per job; stress case at +90% DRAMMeasuredRisk
4Which region, which hour, which grid factor?kgCO2/kWh by region (SG 0.402; Peninsular Malaysia 0.740); hourly placementRegulator printRisk / ESG
5Who can stop the agent at 2 a.m., and does “stop” flatten the book?One name; the last denied action; kill switch vs wind-downDemonstrated, not describedInvestment committee
6Which provider carries your research, and when did you last run the fallback?Share of runs on the top provider; fallback test date and resultTested, with a dateIC / Risk

Two of these on the next quarterly call. Mental model: revealed preference. A manager who cannot fill the cell in 90 days has told you something about operational maturity, not about virtue.

What did Philip and Jared add?

Philip’s section deserves its own post, and will get one. The line I would tattoo on every pitch deck: AI is a component, not a strategy. His best results in the book came from AI correcting a strategy’s decisions rather than generating them — the Corrective AI forex example moved the Sharpe from 0.88 to 1.29 with the annual return barely changing; what halved was the drawdown. And his sharpest warning was about the part that makes headlines: signal generation is where AI pays least and fades fastest.

Jared ran the book’s multi-agent research loop live. Before his segment I asked the audience to hold one assumed number — 40 model calls × 3.91 Wh ≈ 156 Wh per agent research loop, so 100 loops a night ≈ 15.6 kWh, about S$4.40 of electricity and roughly 6.3 kg of CO2 on Singapore’s grid — and to watch whether his trace produced a real one. At desk scale, the kilowatt-hours are cents. The exposure is the blank row.

How was the deck built? With AI — and that is the point

One disclosure I made on the call and will repeat here, because it is a live demonstration of the S and G columns: this deck was expanded with AI, from my prompts, and the message is exactly the one I set out to make.

The thesis, the three lenses, the Receipt–Person–Stop framing, the six questions, the decision on which numbers were binding and which were merely proposals — those were mine, written as a sequence of prompts over a number of evenings. The model’s job was to widen each argument, hunt for the supporting figure, draft the table, and tell me where a claim was thinner than I thought. Every number then went through the same evidence rule the deck preaches: source named, date checked, label attached — measured, estimated, derived or not disclosed. Several claims did not survive that pass and were cut. One that did survive got a footnote saying “derived, not a regulator’s print,” because that is what it is.

There is a fashionable word for AI-assisted output, and it is a farmyard word. The idea is that anything a model touched must be a grey paste of filler with the nutritional value of a press release. I understand where it comes from — the internet has certainly been fed a great deal of it lately. But notice what the word actually describes. Slop is what comes out of a trough when somebody poured slop in. Nobody blames the bucket.

AI is an amplifier, not an author. It amplifies the quality of the prompt and fills the gaps around it. Point it at a vague request and it will confidently produce three hundred words of nothing, beautifully formatted. Point it at a precise thesis with named sources, a stated evidence rule and a person who will delete anything that fails it, and it produces a sixteen-slide argument that a room of CFA charterholders could not find a hole in for fifty minutes. The difference was not the model. It was the person holding it.

Which is precisely the Person column. The human minutes in this deck were not spent typing bullet points; they were spent deciding what the bullet points had to prove, and killing the ones that did not. If a desk’s AI output is farmyard-grade, the ledger already tells you where to look — not at the model’s watt-hours, but at the blank where a human’s minutes should have been.

What I think happens next

Three dated calls, on the record, in addition to the three I made before the event:

  1. By 30 June 2027, at least one frontier model provider publishes a measured watt-hours-per-query figure for a paid API tier, because a large enterprise customer makes it a procurement condition. “Not disclosed” stops being a stable equilibrium once one provider breaks ranks.
  2. By the end of 2027, “human minutes per accepted decision” or a close equivalent appears in a published operational due-diligence questionnaire from a top-tier allocator or ODD consultancy.
  3. Within 12 months, a regulator in a major market publicly asks a supervised firm to demonstrate — not describe — a rehearsed stop of an autonomous research or trading agent.

If I’m wrong, I’ll link back here and say so.

Alternative Perspectives

Contrarian view: The whole “E” argument collapses to FinOps. If halving RAM is the answer, call it cost control and drop the ESG vocabulary, which only invites greenwashing accusations. I half agree — that is exactly why I insist FinOps and ESG are cousins, not twins. The levers that cut both dollars and kWh are the easy 60%. The other 40% — cheaper-but-dirtier regions, bigger models for a headline Sharpe, real-time inference for a daily signal — are where the two diverge, and a DDQ that only measures cost is diligencing half the manager.

Emerging angle: As agents take over more of the loop, S and G merge. When no human is in the loop to count the trials, the “person” blank and the “stop” blank become the same blank — and the Deflated Sharpe becomes a governance control rather than a statistics footnote.

The full deck

The book behind the session is Hands-On AI Trading with Python, QuantConnect, and AWS (Wiley, 2025) by Jiri Pik, Ernest Chan, Jared Broad, Philip Sun and Vivek Singh (Amazon). All code is open source at github.com/QuantConnect/HandsOnAITradingBook.

FAQ

So does AI trading have an ESG problem or not?

Yes — three of them, if you use ESG as lenses rather than a label. E: a production input (watt-hours per decision) with no receipt, in a year when its price doubled. S: the agent does the work the junior used to learn from, so nobody is competent when it is wrong. G: no named person who can stop it, running on the same two providers as everyone else. None is moral. All are P&L.

What is the AI Sustainability Ledger?

A three-column, one-rule framework I presented at the CFA Institute webinar. Columns: Receipt (E), Person (S), Stop (G). Rule: model + version, data known at decision time, every trial counted, point-in-time for price, ESG and the model itself. Example: under Receipt, “kWh per accepted decision, labelled measured / estimated / not disclosed.”

What is the difference between the ESG lens and the ESG label?

The label is a claim — “sustainable”, “AI-powered”, “responsible”. The lens is a question aimed at a risk: what inputs am I short, whose skill am I eroding, what can I stop. Labels can be washed. Lenses produce numbers a manager either has or does not.

Isn’t a trading desk’s compute footprint negligible?

At desk scale, yes — the kilowatt-hours are cents. The exposure is not the emissions; it is holding an unpriced, undisclosed production input while its price doubles, and doing so on a grid that is about to make region and hour reportable.

Where do I start?

Ask two of the six questions on your next manager call and record every “not disclosed”. Inside your own shop, create the one number that does not exist yet: human minutes per accepted decision.


Disclaimer: This reflects my personal views and experience, not financial advice. Past performance doesn’t guarantee future results. Performance figures quoted are illustrative examples from the book, not live trading results. Energy figures marked as assumed or derived are my own estimates, not measurements.

Jiri Pik is the founder of RocketEdge, an AI fintech company based in Singapore, and a co-author of Hands-On AI Trading with Python, QuantConnect, and AWS (Wiley). Follow him on LinkedIn and X for more.

keyboard_arrow_up