I recently joined David Greely on the Smarter Markets podcast for an episode called 'How to Raise Your Agent'. We covered four things that shape how I think about the next few years: chat as the interface that won, chat data as an untapped resource, AI as the tool that finally makes that resource usable, and governance as the layer that decides who gets to act on it safely.
The full episode is on Smarter Markets Media and on Smarter Markets' YouTube channel. Here are my key takeaways from the conversation.
I opened by tracing how people interact with market data: GUIs have been around since the '90s, letting people point and click on a screen. That evolved into APIs, where systems could do the pointing and clicking instead. Agents are the next step.
None of that changes the fact that 75% of OTC trading still happens on chat, voice and email. What changes is what sits on top of the conversation.
I compared chat data to Venezuela's oil reserves on the podcast: the largest in the world, unusable because extraction costs too much. Chat has worked the same way in financial and commodities markets. Every price, RFQ and confirmation that moves through a trading desk's chat window is real information. Almost none of it gets captured, because turning free-text conversation into structured data has been too expensive to do at scale.
Large language models change that cost equation. It's the argument behind our Chat Data Mining capability at ipushpull: the same tooling that captures a trade confirmation can be pointed at any market, and accuracy improves as more messages get validated. I talked about Koch Energy Services on the show, who reached 90%+ field-level accuracy on physical gas trade capture after building up a library of their own desk-specific syntax. Read the Koch Energy Services case study.
We built the autonomy slider because we kept having the same conversation with clients and compliance teams: how far should AI be allowed to run in a workflow before a person needs to see the output? It's the framework I now use with clients to answer that question, and I trace the underlying idea back to Andrej Karpathy's work and to language now appearing in regulatory guidance from IOSCO.
It runs across four levels. Human-in-control: a person decides everything, AI plays no active role. Human-in-the-loop: AI acts, and a person approves at each gate before anything moves forward. Human-on-the-loop: AI runs the workflow while a person monitors it, stepping in if something looks wrong. Human-out-of-the-loop: AI acts independently, without a person reviewing it in real time.
In a large regulated financial institution, the only place you'll currently see anything close to human-out-of-the-loop is around post-trade processes, because the trade's already been made and the risk of something going wrong is lower.
Where a given workflow sits depends on the money at stake, the accuracy the task demands, and how much oversight is realistic to expect from a person. You can move the slider yourself and see what each level means for compliance and governance on our autonomy slider page (and for a deeper dive, read my essay).
IOSCO has since published its own supervisory toolkit for AI in capital markets, built around these same four levels. I go deeper on what that means for regulated trading firms in What IOSCO's new AI supervisory toolkit means for regulated trading firms.
My argument against treating governance as compliance overhead: think of it like an F1 car. You're constantly fine-tuning it, and the better it is, the faster and better the performance you get out of it. If you've built the right controls in from the start, you don't need to test every new tool from scratch. You can move through the autonomy slider with a record you can show a regulator.
One of my favourite areas of innovation is around RFQ negotiations. A trader sends a request to several dealers, picks the best price, and the rest of the responses disappear. I described an agent that would instead flag which dealer was fastest, which usually prices best when they do respond, and which pattern is worth acting on next time. That data was always there. It just went nowhere. We wrote about this in more detail in Your RFQ responses already rank your dealers. No one's looked.
I ended the episode on agent-to-agent trading. By the end of the decade, agents will be trading with other agents, and the industry doesn't yet have the identity, permissioning and audit layers to make that safe. That's exciting. It's also a little bit scary if you haven't got the governance layers in place first.
Listen to the full episode for the rest of the conversation, including how our chat data mining tooling takes a market's accuracy from around 50-65% out of the box to 90%+ with a small amount of human validation.