Recherche & preuves
Explorez les travaux scientifiques dont s'inspire AlphaIntel. Ces articles sont évalués par des pairs ; notre mise en œuvre, elle, est la nôtre, et nous la décrivons ouvertement.
Les articles sont présentés dans leur langue de publication, l'anglais. Les titres, résumés et mots-clés ne sont pas traduits : un résumé scientifique traduit cesse d'être citable tel quel.
Glissez pour voir tous les filtres
Advanced Financial Reasoning at Scale: LLMs on CFA Level III
A comprehensive evaluation of state-of-the-art LLMs on the CFA Level III exam. OpenAI o4-mini achieved a score of 79.1%, and Gemini 2.5 Flash reached 77.3%, demonstrating expert-level financial reasoning capabilities.
Ce que cela signifie pour vous
Les modèles d'IA récents passent l'examen CFA niveau III, la certification de référence des analystes professionnels.
Sentiment Trading with Large Language Models
This study compares dictionary-based methods with modern LLMs for sentiment analysis. Strategies based on OPT-66B generated a Sharpe Ratio of 3.05, significantly outperforming traditional methods (Sharpe 1.23).
Ce que cela signifie pour vous
Lire le sentiment avec un LLM donne un ratio de Sharpe plus de deux fois supérieur aux méthodes de dictionnaire classiques.
QuantAgents: Towards Multi-agent Financial System via Simulated Trading
Presents QuantAgent, a multi-agent system that divides trading into specialized roles (Indicator, Pattern, Trend, Risk). Achieved 111.87% annualized return and a Sharpe Ratio of 2.02 in backtesting.
Ce que cela signifie pour vous
Répartir l'analyse entre des agents spécialisés (indicateurs, figures, tendance, risque) donne de bien meilleurs rendements ajustés du risque qu'un seul modèle généraliste.
FinChain: A Symbolic Benchmark for Verifiable Chain-of-Thought Financial Reasoning
Introduces ChainEval to measure the quality of financial reasoning. Shows that Chain-of-Thought (CoT) prompting significantly reduces logic errors and that reasoning models correlate strongly with expert human judgment.
Ce que cela signifie pour vous
Un modèle qui détaille son raisonnement étape par étape commet nettement moins d'erreurs de logique en analyse financière.
TradingAgents: Multi-Agents LLM Financial Trading Framework
Explores the debate mechanism between opposing agents. The study finds that agent debate reduces hallucinations and improves risk-adjusted returns (Sortino/Sharpe ratios) compared to single-agent models.
Ce que cela signifie pour vous
Faire débattre des agents entre eux, et les obliger à se contredire, réduit les erreurs et améliore le rendement ajusté du risque par rapport à un modèle unique.
Single-agent or Multi-agent Systems? Why Not Both?
Comparative analysis showing that Multi-Agent Systems (MAS) offer superior accuracy for complex tasks. A hybrid architecture can improve precision by 1.1% to 12% while optimizing inference costs.
Ce que cela signifie pour vous
Sur une décision financière complexe, un système multi-agents est mesurablement plus juste qu'un modèle seul.
Deep Reinforcement Learning for Automated Stock Trading
Demonstrates that an ensemble of RL algorithms (PPO, A2C, DDPG) adapts better to market regime changes than individual algorithms, generating superior Sharpe ratios on the Dow Jones index.
Ce que cela signifie pour vous
Combiner plusieurs algorithmes d'apprentissage par renforcement encaisse mieux les changements de régime de marché que n'importe lequel pris seul.
Optimal Profit-Making Strategies with Algorithmic Trading
A longitudinal study (2006-2023) on the CSI 300 index showing that Support Vector Machines (SVM) generated an excess return of 60.52%, proving the long-term robustness of classical ML methods.
Ce que cela signifie pour vous
Les méthodes d'apprentissage classiques, comme les SVM, tiennent la distance : 17 ans de données de marché le montrent.
Reinforcement Learning for Deep Portfolio Optimization
Integrates Modern Portfolio Theory constraints directly into the RL reward function (Deep Portfolio Optimization). Maximizes portfolio value while strictly adhering to risk constraints.
Ce que cela signifie pour vous
Inscrire les contraintes de risque directement dans la fonction de récompense donne des portefeuilles qui cherchent le rendement sans dépasser la perte maximale acceptée.
Modeling News Interactions and Influence for Financial Market Prediction
Proves that fusing textual data (news) with price action (FININ model) increases the daily Sharpe Ratio by +0.429 compared to using price data alone.
Ce que cela signifie pour vous
Croiser le texte des actualités avec les prix améliore sensiblement la qualité des prévisions par rapport aux prix seuls.
Benchmarking LLMs for Target-Based Financial Sentiment Analysis
Research indicating that generative models (like GPT-4, DeepSeek) now outperform specialized older models (FinBERT) in zero-shot sentiment analysis tasks.
Ce que cela signifie pour vous
Les LLM généralistes lisent désormais le sentiment de marché mieux que les modèles spécialisés en finance, sans réglage supplémentaire.
Dynamic Stop Loss Strategy with Deep Reinforcement Learning
Shows that RL agents can learn optimal dynamic stop-loss policies that adapt to market volatility, significantly improving PnL and reducing maximum drawdown compared to static rules.
Ce que cela signifie pour vous
Un stop de protection qui s'adapte à la volatilité fait mieux qu'un pourcentage fixe, sur le gain comme sur la perte maximale.
FPGA Acceleration for Financial Machine Learning
Validates the use of FPGA accelerators to achieve millisecond-level inference for complex ML models, maintaining >90% accuracy while enabling high-frequency execution.
Ce que cela signifie pour vous
Accélérer un modèle par le matériel permet de le faire tourner en moins d'une milliseconde : c'est l'infrastructure du trading haute fréquence institutionnel, pas celle d'un investisseur particulier. Nous la citons parce qu'elle marque la frontière que nous ne franchissons pas : AlphaIntel travaille sur des données journalières et vous laisse décider, là où cette littérature vise la milliseconde et l'exécution automatique.
Alpha-GPT: Human-AI Interactive Alpha Mining for Quantitative Investment
Introduces a new alpha mining paradigm by introducing human-AI interaction and a novel prompt engineering algorithmic framework leveraging large language models. Alpha-GPT provides a heuristic way to understand quant researchers' ideas and outputs creative, insightful, and effective alphas.
Ce que cela signifie pour vous
Associer l'expertise humaine à la créativité d'un LLM produit de meilleurs signaux que l'une ou l'autre approche seule.
LLMFactor: Extracting Profitable Factors through Prompts for Explainable Stock Movement Prediction
Introduces LLMFactor, a novel framework employing Sequential Knowledge-Guided Prompting (SKGP) to identify factors influencing stock movements using LLMs. Extracts factors more directly related to stock market dynamics, providing clear explanations for complex temporal changes.
Ce que cela signifie pour vous
Un LLM sait extraire les facteurs précis qui expliquent le mouvement d'un titre, et dire pourquoi, au lieu de sortir un signal nu.
HedgeAgents: A Balanced-aware Multi-agent Financial Trading System
Introduces HedgeAgents, an innovative multi-agent system aimed at bolstering system robustness via hedging strategies. The framework features a central fund manager and multiple hedging experts, achieving 70% annualized return and 400% total return over 3 years.
Ce que cela signifie pour vous
Un agent gérant de fonds qui coordonne des agents de couverture spécialisés a obtenu 70 % de rendement annualisé sur trois ans, en conditions de test.
LLMs for Quantitative Investment Research: A Practitioner's Guide
A practitioner-oriented review (UCL / DWS) of how LLMs are reshaping quantitative investment research across three fronts: research assistance, text-based signal extraction, and systemising expert judgment. Documents the field's hard empirical limits (temporal leakage, memorisation, behavioural biases, reproducibility) and provides governance guidelines plus an evaluation checklist for deploying LLMs in production research pipelines.
Ce que cela signifie pour vous
Le consensus du secteur après trois ans de déploiement : un LLM apporte une vraie valeur comme extracteur et interprète de signaux à l'intérieur d'une chaîne déterministe encadrée, pas comme moteur de prévision laissé seul.
F²Agent: Modality-Aware Fusion of Specialist Agents for Financial Trading
Most LLM trading systems merge market data, indicators, news and sentiment by concatenating them into a single prompt, which lets the textual signal dominate and loses cross-modal dependencies. F²Agent instead assigns one specialist encoder per modality and fuses them through learned modality-aware attention, regularised for prior diversity and for stability under single-modality perturbation. Across six assets it ranks first on annualized return against sixteen baselines, and its ablation shows the fusion layer, not the agents, carries the gain: replacing it with plain concatenation drops annualized return from 50.1% to 19.2% on AAPL.
Ce que cela signifie pour vous
Le poids de chaque source de données dans un verdict devrait être un nombre explicite et vérifiable, pas une information noyée dans un bloc de texte.
Calibration-Induced Degeneracy: When an Expensive LLM Feature Receives Zero Weight
A fully audit-trailed case study in which an LLM scores the market importance of daily news headlines, and that score is added to implied-volatility baselines to forecast next-day equity risk. Calibration assigned the feature a weight of exactly zero, making every augmented forecast mathematically identical to its baseline, after the full-history inference had already been purchased. A near-free control that merely counts headlines did improve variance forecasting. The paper prescribes a viability checkpoint, fit first, perturb second, acquire last, that detects a feature which cannot move the output before any inference is paid for.
Ce que cela signifie pour vous
Avant de payer pour un signal sophistiqué, vérifiez deux choses : qu'il peut mécaniquement changer la réponse, et qu'il bat une alternative quasi gratuite.
MemArbiter: Arbitrating Memory at the Moment of Decision
Identifies the Memory-Action Gap: in long-horizon agents the bottleneck is neither storing information nor retrieving it, but arbitrating which of it surfaces at the moment a decision is made. MemArbiter organises memory into functional banks (goal, task state, constraint, episodic, reference), scores decision relevance per bank and per item, and applies a temporal gate that deliberately shields goals and constraints from time decay. On unseen ALFWorld tasks it reaches 82.8% success at a fixed 500-token memory budget against 61.9% for flat retrieval, and halves the rate at which an agent repeats an action that has already failed.
Ce que cela signifie pour vous
Une contrainte que le système connaît mais ne présente pas au modèle au moment de décider est, en pratique, une contrainte qui n'existe pas.
Harness the Memory: A Holistic Benchmark of Agent Memory Substrates
Evaluates eleven memory substrates across seven families, three backbones and four task suites under twenty-six metrics. No substrate dominates: the winner inverts by regime, with structured graphs leading on dialogue question answering and cheap flat retrieval leading on code. Widening retrieval helps question answering but measurably degrades sequential decision-making, with attention probes showing mass draining from the action context into the retrieved block. Production-grade memory systems are also shown to cost between 2,700 and 9,000 auxiliary LLM calls per long history, an overhead usually reported nowhere.
Ce que cela signifie pour vous
Récupérer plus de contexte n'est pas automatiquement mieux : passé un certain point, cela chasse l'information dont le modèle a réellement besoin pour agir.
Stealing Reasoning Traces from Proprietary LLM APIs
Reasoning models return their chain-of-thought to the client as an opaque encrypted block that the client must replay on later calls. These blocks are shown to be interchangeable across sessions, users and sibling models, so a weaker model from the same provider can be coerced into transcribing a stronger model's hidden reasoning verbatim. Decoding 315,320 blocks scraped from publicly shared sessions recovered 704 distinct real secrets, including 62 API keys and 33 passwords, 64 of which appeared nowhere in the visible transcript, making plaintext-only sanitization ineffective. The displayed reasoning summary is also measured to hide roughly five times more content than it shows.
Ce que cela signifie pour vous
Le raisonnement caché d'un modèle contient régulièrement des éléments sensibles absents de sa réponse visible : il doit être traité comme une donnée à protéger, pas comme un artefact inoffensif.
Model Discovery Agent: LLM-Assisted Bayesian Experiment Design
Couples an LLM used strictly as a proposer of candidate mechanisms with standard Bayesian machinery: sequential Monte Carlo for posteriors and evidence, and value-of-information to choose the next experiment. Across physics, chemistry and neuroscience benchmarks it recovers the correct mechanism far more data-efficiently than an LLM agent working alone. The most transferable result is a robustness one: because the decision sits in the deterministic layer, accuracy stays between 89% and 94% regardless of which model proposes, whereas the pure LLM agent swings from 26% to 81% depending on the model behind it.
Ce que cela signifie pour vous
Quand la décision finale repose sur une mécanique déterministe plutôt que sur le modèle lui-même, changer de modèle cesse d'être une source de comportements imprévisibles.
Why Better Models Can Create Riskier Systems: Evidence from LLM Agents in Financial Markets
Argues that as language models grow more capable, their errors become more correlated with one another, so a market populated by many LLM agents can carry a risk floor that adding more agents never diversifies away. Across the models studied, the correlation of residual decisions between pairs of models rises with capability, while sharing a provider shows no significant effect. In a simulated single-asset market, more LLM traders improve price discovery under normal conditions, but when every agent reads the same misleading commentary, tracking error rises well above a noise-trader baseline for two of the three model families tested. The authors also show that mixing model families removes only the family-specific part of the correlation; the part that comes from a shared information environment remains.
Ce que cela signifie pour vous
Des analystes qui tournent sur le même modèle et lisent les mêmes news pèsent moins que des avis vraiment distincts, car la vraie diversité vient d'informations indépendantes et pas seulement de rôles différents.
MemRiskBench: Trace-Aware Risk-Preserving Evaluation for Long-Horizon LLM Agents
Argues that average scores hide the rare but serious failures of agents with long-term memory. Defines five risk types (stale facts after an update, unresolved conflicting facts, leakage across users, reuse of revoked memories, and gradual decay of standing constraints) and checks each one deterministically against the agent's execution trace instead of relying on an LLM judge. Across 120 scripted episodes and five small open models, overall pass rates conceal large per-risk gaps, with cross-user leakage the weakest category for every model tested. A subset that keeps only the episodes where models disagree reproduces the full ranking at a fifth of the cost.
Ce que cela signifie pour vous
Un assistant qui se souvient de vous doit aussi savoir quand une information est périmée, et l'oublier quand vous le demandez, ce qu'une bonne note moyenne ne garantit pas.
Recursive Self-Improvement of AI Research Agents
Uses one AI agent to rewrite the code of another research agent in a two-level loop: the inner agent optimizes each task against a public score, while the outer loop keeps a rewrite only if it improves a private, held-out score the inner agent never sees. Over an eight-day autonomous run the system accepted seven improvements, including a bandit over drafting strategies and bounded summaries of long logs, and the resulting agent matched or beat a human-designed baseline on four held-out benchmarks. Its measured rate of reward hacking on a kernel-optimization test fell from 55% to 32%, on a small sample reported without error bars.
Ce que cela signifie pour vous
Un système réglé et noté sur les mêmes exemples finit par apprendre l'examen, et seule une série de cas tenue à l'écart de chaque réglage donne une note honnête.