Research & Evidence
Explore the research AlphaIntel draws on. These papers are peer-reviewed; our implementation of them is ours, and we describe it openly.
Papers are shown in their language of publication, English. Titles, abstracts and keywords are not translated: a translated scientific abstract stops being quotable as-is.
Swipe to see all filters
Advanced Financial Reasoning at Scale: LLMs on CFA Level III
A comprehensive evaluation of state-of-the-art LLMs on the CFA Level III exam. OpenAI o4-mini achieved a score of 79.1%, and Gemini 2.5 Flash reached 77.3%, demonstrating expert-level financial reasoning capabilities.
What this means for traders
Modern AI models now pass the CFA Level III exam, the same benchmark used to certify professional analysts.
Sentiment Trading with Large Language Models
This study compares dictionary-based methods with modern LLMs for sentiment analysis. Strategies based on OPT-66B generated a Sharpe Ratio of 3.05, significantly outperforming traditional methods (Sharpe 1.23).
What this means for traders
LLM-based sentiment reading produces Sharpe ratios more than twice as high as traditional dictionary methods.
QuantAgents: Towards Multi-agent Financial System via Simulated Trading
Presents QuantAgent, a multi-agent system that divides trading into specialized roles (Indicator, Pattern, Trend, Risk). Achieved 111.87% annualized return and a Sharpe Ratio of 2.02 in backtesting.
What this means for traders
Dividing analysis into specialist agents (indicator, pattern, trend, risk) produces significantly better risk-adjusted returns than a single generalist model.
FinChain: A Symbolic Benchmark for Verifiable Chain-of-Thought Financial Reasoning
Introduces ChainEval to measure the quality of financial reasoning. Shows that Chain-of-Thought (CoT) prompting significantly reduces logic errors and that reasoning models correlate strongly with expert human judgment.
What this means for traders
AI models that show their reasoning step-by-step make far fewer logical errors in financial analysis.
TradingAgents: Multi-Agents LLM Financial Trading Framework
Explores the debate mechanism between opposing agents. The study finds that agent debate reduces hallucinations and improves risk-adjusted returns (Sortino/Sharpe ratios) compared to single-agent models.
What this means for traders
Having AI agents debate and challenge each other reduces errors and improves risk-adjusted returns compared to using a single AI model.
Single-agent or Multi-agent Systems? Why Not Both?
Comparative analysis showing that Multi-Agent Systems (MAS) offer superior accuracy for complex tasks. A hybrid architecture can improve precision by 1.1% to 12% while optimizing inference costs.
What this means for traders
For complex financial decisions, multi-agent systems are measurably more accurate than a single AI model.
Deep Reinforcement Learning for Automated Stock Trading
Demonstrates that an ensemble of RL algorithms (PPO, A2C, DDPG) adapts better to market regime changes than individual algorithms, generating superior Sharpe ratios on the Dow Jones index.
What this means for traders
Combining multiple reinforcement learning algorithms handles market regime changes better than any single algorithm.
Optimal Profit-Making Strategies with Algorithmic Trading
A longitudinal study (2006-2023) on the CSI 300 index showing that Support Vector Machines (SVM) generated an excess return of 60.52%, proving the long-term robustness of classical ML methods.
What this means for traders
Classical ML methods like SVMs have proven durable alpha-generators over 17 years of market data.
Reinforcement Learning for Deep Portfolio Optimization
Integrates Modern Portfolio Theory constraints directly into the RL reward function (Deep Portfolio Optimization). Maximizes portfolio value while strictly adhering to risk constraints.
What this means for traders
Baking risk constraints directly into the AI reward function produces portfolios that maximize returns while respecting hard drawdown limits.
Modeling News Interactions and Influence for Financial Market Prediction
Proves that fusing textual data (news) with price action (FININ model) increases the daily Sharpe Ratio by +0.429 compared to using price data alone.
What this means for traders
Combining news text with price data meaningfully improves prediction quality over price-only models.
Benchmarking LLMs for Target-Based Financial Sentiment Analysis
Research indicating that generative models (like GPT-4, DeepSeek) now outperform specialized older models (FinBERT) in zero-shot sentiment analysis tasks.
What this means for traders
General-purpose LLMs now outperform finance-specific models at reading market sentiment, with no fine-tuning needed.
Dynamic Stop Loss Strategy with Deep Reinforcement Learning
Shows that RL agents can learn optimal dynamic stop-loss policies that adapt to market volatility, significantly improving PnL and reducing maximum drawdown compared to static rules.
What this means for traders
AI-driven stop-losses that adapt to volatility outperform fixed percentage rules in both profit and drawdown control.
FPGA Acceleration for Financial Machine Learning
Validates the use of FPGA accelerators to achieve millisecond-level inference for complex ML models, maintaining >90% accuracy while enabling high-frequency execution.
What this means for traders
Hardware-accelerated ML can run complex models at sub-millisecond speed: the infrastructure behind institutional HFT, not a private investor's. We cite it because it marks the line we do not cross: AlphaIntel works on daily data and leaves the decision to you, where this literature targets the millisecond and automated execution.
Alpha-GPT: Human-AI Interactive Alpha Mining for Quantitative Investment
Introduces a new alpha mining paradigm by introducing human-AI interaction and a novel prompt engineering algorithmic framework leveraging large language models. Alpha-GPT provides a heuristic way to understand quant researchers' ideas and outputs creative, insightful, and effective alphas.
What this means for traders
Combining human domain knowledge with LLM creativity produces better alpha signals than either approach alone.
LLMFactor: Extracting Profitable Factors through Prompts for Explainable Stock Movement Prediction
Introduces LLMFactor, a novel framework employing Sequential Knowledge-Guided Prompting (SKGP) to identify factors influencing stock movements using LLMs. Extracts factors more directly related to stock market dynamics, providing clear explanations for complex temporal changes.
What this means for traders
LLMs can extract the specific factors driving a stock move and explain why, not just output a signal.
HedgeAgents: A Balanced-aware Multi-agent Financial Trading System
Introduces HedgeAgents, an innovative multi-agent system aimed at bolstering system robustness via hedging strategies. The framework features a central fund manager and multiple hedging experts, achieving 70% annualized return and 400% total return over 3 years.
What this means for traders
A fund-manager agent coordinating specialist hedging agents achieved 70% annualized return over three years.
LLMs for Quantitative Investment Research: A Practitioner's Guide
A practitioner-oriented review (UCL / DWS) of how LLMs are reshaping quantitative investment research across three fronts: research assistance, text-based signal extraction, and systemising expert judgment. Documents the field's hard empirical limits (temporal leakage, memorisation, behavioural biases, reproducibility) and provides governance guidelines plus an evaluation checklist for deploying LLMs in production research pipelines.
What this means for traders
Industry consensus after three years of deployment: LLMs add real value as signal extractors and interpreters inside governed, deterministic pipelines, not as standalone forecasting engines.
F²Agent: Modality-Aware Fusion of Specialist Agents for Financial Trading
Most LLM trading systems merge market data, indicators, news and sentiment by concatenating them into a single prompt, which lets the textual signal dominate and loses cross-modal dependencies. F²Agent instead assigns one specialist encoder per modality and fuses them through learned modality-aware attention, regularised for prior diversity and for stability under single-modality perturbation. Across six assets it ranks first on annualized return against sixteen baselines, and its ablation shows the fusion layer, not the agents, carries the gain: replacing it with plain concatenation drops annualized return from 50.1% to 19.2% on AAPL.
What this means for traders
How much each data source counts toward a verdict should be an explicit, inspectable number, not something buried inside a block of text.
Calibration-Induced Degeneracy: When an Expensive LLM Feature Receives Zero Weight
A fully audit-trailed case study in which an LLM scores the market importance of daily news headlines, and that score is added to implied-volatility baselines to forecast next-day equity risk. Calibration assigned the feature a weight of exactly zero, making every augmented forecast mathematically identical to its baseline, after the full-history inference had already been purchased. A near-free control that merely counts headlines did improve variance forecasting. The paper prescribes a viability checkpoint, fit first, perturb second, acquire last, that detects a feature which cannot move the output before any inference is paid for.
What this means for traders
Before paying for a sophisticated signal, check two things: that it can mechanically change the answer at all, and that it beats a trivially cheap alternative.
MemArbiter: Arbitrating Memory at the Moment of Decision
Identifies the Memory-Action Gap: in long-horizon agents the bottleneck is neither storing information nor retrieving it, but arbitrating which of it surfaces at the moment a decision is made. MemArbiter organises memory into functional banks (goal, task state, constraint, episodic, reference), scores decision relevance per bank and per item, and applies a temporal gate that deliberately shields goals and constraints from time decay. On unseen ALFWorld tasks it reaches 82.8% success at a fixed 500-token memory budget against 61.9% for flat retrieval, and halves the rate at which an agent repeats an action that has already failed.
What this means for traders
A constraint the system knows about but does not put in front of the model at decision time is, in practice, a constraint that does not exist.
Harness the Memory: A Holistic Benchmark of Agent Memory Substrates
Evaluates eleven memory substrates across seven families, three backbones and four task suites under twenty-six metrics. No substrate dominates: the winner inverts by regime, with structured graphs leading on dialogue question answering and cheap flat retrieval leading on code. Widening retrieval helps question answering but measurably degrades sequential decision-making, with attention probes showing mass draining from the action context into the retrieved block. Production-grade memory systems are also shown to cost between 2,700 and 9,000 auxiliary LLM calls per long history, an overhead usually reported nowhere.
What this means for traders
Retrieving more context is not automatically better: past a point it crowds out the information the model actually needs to act on.
Stealing Reasoning Traces from Proprietary LLM APIs
Reasoning models return their chain-of-thought to the client as an opaque encrypted block that the client must replay on later calls. These blocks are shown to be interchangeable across sessions, users and sibling models, so a weaker model from the same provider can be coerced into transcribing a stronger model's hidden reasoning verbatim. Decoding 315,320 blocks scraped from publicly shared sessions recovered 704 distinct real secrets, including 62 API keys and 33 passwords, 64 of which appeared nowhere in the visible transcript, making plaintext-only sanitization ineffective. The displayed reasoning summary is also measured to hide roughly five times more content than it shows.
What this means for traders
A model's hidden reasoning routinely contains sensitive material that never appears in its visible answer, so it must be treated as data worth protecting rather than as a harmless artifact.
Model Discovery Agent: LLM-Assisted Bayesian Experiment Design
Couples an LLM used strictly as a proposer of candidate mechanisms with standard Bayesian machinery: sequential Monte Carlo for posteriors and evidence, and value-of-information to choose the next experiment. Across physics, chemistry and neuroscience benchmarks it recovers the correct mechanism far more data-efficiently than an LLM agent working alone. The most transferable result is a robustness one: because the decision sits in the deterministic layer, accuracy stays between 89% and 94% regardless of which model proposes, whereas the pure LLM agent swings from 26% to 81% depending on the model behind it.
What this means for traders
When the final decision rests on deterministic machinery rather than on the model itself, swapping the underlying model stops being a source of unpredictable behavior.
Why Better Models Can Create Riskier Systems: Evidence from LLM Agents in Financial Markets
Argues that as language models grow more capable, their errors become more correlated with one another, so a market populated by many LLM agents can carry a risk floor that adding more agents never diversifies away. Across the models studied, the correlation of residual decisions between pairs of models rises with capability, while sharing a provider shows no significant effect. In a simulated single-asset market, more LLM traders improve price discovery under normal conditions, but when every agent reads the same misleading commentary, tracking error rises well above a noise-trader baseline for two of the three model families tested. The authors also show that mixing model families removes only the family-specific part of the correlation; the part that comes from a shared information environment remains.
What this means for traders
Several analysts running on the same model and reading the same news are closer to a single opinion than to many: real diversity comes from independent information, not just from different personas.
MemRiskBench: Trace-Aware Risk-Preserving Evaluation for Long-Horizon LLM Agents
Argues that average scores hide the rare but serious failures of agents with long-term memory. Defines five risk types (stale facts after an update, unresolved conflicting facts, leakage across users, reuse of revoked memories, and gradual decay of standing constraints) and checks each one deterministically against the agent's execution trace instead of relying on an LLM judge. Across 120 scripted episodes and five small open models, overall pass rates conceal large per-risk gaps, with cross-user leakage the weakest category for every model tested. A subset that keeps only the episodes where models disagree reproduces the full ranking at a fifth of the cost.
What this means for traders
An assistant that remembers you should also know when a fact is out of date and forget it when you ask; a good average score says nothing about either.
Recursive Self-Improvement of AI Research Agents
Uses one AI agent to rewrite the code of another research agent in a two-level loop: the inner agent optimizes each task against a public score, while the outer loop keeps a rewrite only if it improves a private, held-out score the inner agent never sees. Over an eight-day autonomous run the system accepted seven improvements, including a bandit over drafting strategies and bounded summaries of long logs, and the resulting agent matched or beat a human-designed baseline on four held-out benchmarks. Its measured rate of reward hacking on a kernel-optimization test fell from 55% to 32%, on a small sample reported without error bars.
What this means for traders
A system tuned and graded on the same examples ends up learning the exam; the only honest score comes from cases held back from every adjustment.