The "Buy to Build" Paradigm: Redefining Data Strategy for Buy-Side Execution
First published by TabbFORUM
Rob Laible, Head of Americas, BMLL
Historically, quantitative trading desks operated with highly ad hoc workflows, spending immense amounts of time responding to internal queries and curating raw market data. Rather than focusing solely on alpha generation, the quantitative trading landscape required large teams of IT personnel and data engineers merely to stitch, clean, and normalize raw feeds. However, the modern buy-side desk has fundamentally shifted to focus entirely on deploying advanced analytics and optimizing execution.
At the recent BMLL Client Summit in New York, leading buy-side practitioners gathered to discuss this evolution, exploring how cutting-edge data ecosystems and practical AI applications are transforming their trading desks.
Here’s what we heard and learned on the day.
Data Quality and the ‘Buy to Build’ Paradigm
Maintaining in-house databases for streaming tick data historically required enormous engineering effort. Desks had to manage real-time feeds, fix continuous data gaps, update complex symbol watchlists, and navigate restrictive legal agreements with each individual exchange and data provider. One practitioner shared an anecdote highlighting these legacy burdens: after storing streaming global equities data for 10 to 15 years, their vendor was acquired, and new legal teams demanded three times the cost just to read the historical data they had already collected. Furthermore, with legacy processes, if a desk forgot to explicitly tell the system to record a specific symbol, that data was simply lost forever.
Even the most sophisticated, newly built predictive models rely on data for 99% of their foundation; if this underlying data is not pristine, the models operate on a strict "garbage in, garbage out" basis. Because trading desks have an infinite number of questions to answer in a highly constrained time, there has been a strategic shift in resourcing. The modern consensus is increasingly embracing a "buy to build" philosophy: desks should "buy" generalized, clean, and normalized data from third-party vendors, reserving internal engineering capital strictly to "build" highly sophisticated, proprietary products that address their unique organizational nuances.
Accelerating Research with Cloud Data Labs
Transitioning from legacy on-premise infrastructure to cloud-based data labs and compute platforms has dramatically accelerated the speed of research and analytics. The sheer compute power available today allows massive analytical tasks to be processed in a fraction of the time. For example, running fill markouts, which can require evaluating 14 marks on each side of over 100 million fills annually, can now be executed for an entire month using cloud infrastructure in the exact same amount of time it previously took to process a single day's worth of data on legacy on-premise systems.
Furthermore, utilizing normalized third-party data ensures that a firm's research environment and its real-time production models use the exact same foundational reference data. This creates a frictionless workflow, finally eliminating the persistent historical frustration of a trading model performing perfectly in a simulated research lab, only to fail in real-world production due to mismatched reference datasets.
Simplifying Global Market Microstructure
Analyzing global markets requires an intricate understanding of complex, region-specific market structures, such as periodic auctions, systematic internalizers, and off-exchange prints. Building custom orchestration layers and normalization pipelines (ETL layers) for every individual market's specific trade classification codes is an incredibly inefficient use of quantitative resources.
By leveraging pre-enriched datasets like BMLL's Trades Plus, a dataset born directly out of the BMLL Client Product Advisory Board discussions, analysts gain access to a combination of datasets and metrics that are super easy to access and apply directly at the trade level. With these built-in trade classifications, analysts can query global securities and evaluate execution performance seamlessly across regions. This allows traders and analysts to instantly get the context they need to understand the economic activity in the market without having to think about, write, or maintain the underlying normalization code.
Cutting Through AI Hype: LLMs as Force Multipliers
Despite the massive industry hype surrounding Artificial Intelligence, panelists stressed the need to cut through title inflation, where statistics became machine learning, which then became AI. Currently, there is no role for autonomous or agentic AI to make independent execution or trade routing decisions. Trade routing remains firmly governed by traditional supervised and unsupervised machine learning models operating within strict, deterministic boundaries.
However, Large Language Models (LLMs) are acting as an incredibly powerful force multiplier for analysts and developers. In the past, generating trade signals from corporate events required massive teams to hunt through SEC filings, 13F reports, and unstructured data. Today, LLMs can instantly consolidate and summarize this vast amount of unstructured information, allowing a single analyst to accomplish what previously took dozens of developers and engineers, driving rapid, informed decision-making.
Revolutionizing Post-Trade Reporting
AI is also changing the face of post-trade execution reporting. Instead of relying on static, template-based post-trade reports that simply chart execution costs, LLMs are being utilized to analyze the granular, underlying time-series data of every security. The AI can help identify where and why value was created or destroyed during a trade (such as an incorrect call on sector basis risk or legging into a trade improperly) and surface only the most relevant insights on the report's front page.
Crucially, this capability requires active effort from the trading desk. An LLM cannot replicate the contextual understanding of an experienced trader without training and guidance. Desks must actively mentor the AI, much like they would train a junior employee, to ‘think like a trader’ and ‘talk like a trader’. By utilizing Retrieval-Augmented Generation (RAG) in conjunction with human feedback, such as pre and post-corrections, desks can continuously refine the LLM's vocabulary, its analytical focus, and the quality of its output until it learns to effectively communicate complex market dynamics.
A New Era of Execution for the Buy-Side Trading Desk
The convergence of frictionless data, cloud computing power, and targeted AI integration is enabling the buy-side to strip away legacy IT burdens. By shedding the heavy lifting of data curation and infrastructure maintenance, quantitative trading desks are finally free to focus their specialized resources purely on optimizing execution, generating actionable insights, and building the sophisticated trading strategies of the future.