Skip to content

Ayush Mukhi

Finance and business, George Mason University

I work on market microstructure and portfolio allocation. I report what the tests show, including the hypothesis that failed.

GitHub

Predictability exists. It dies at costs.

market-entropy, a pre-registered study

Order flow predicts crypto price moves seconds ahead, consistently and out of sample. The effect is an order of magnitude too small to survive trading fees, and the entropy feature the study was built around added nothing. Both results are reported in full.

Median out-of-sample R² at 1 second is 3.6% for BTC and 2.6% for ETH. It is positive on 97% of 75 held-out days, across 1,020 configurations enumerated in advance and all reported.

Read the research noteSource repository

Median per-day out-of-sample R² by prediction horizon. Walk-forward, monthly refits, held-out days only. Decays toward zero: about 0.2 to 0.3% at 60 seconds.

Gross predictability dies at trading costs

Turning the forecasts into a simple trading rule earns positive gross PnL on most days. Charging any realistic fee erases it. The breakeven one-way cost is 0.17 to 0.33 bps per unit of turnover, against multi-bp taker fees on the venue the data comes from.

c∗  =  ∑dgross PnLd∑dturnoverdc^{*} \;=\; \frac{\sum_{d} \text{gross PnL}_{d}}{\sum_{d} \text{turnover}_{d}}

The flat one-way cost in bps at which net PnL crosses zero, as computed in src/study/costs.py.

Breakeven one-way cost vs the study’s fee assumption. OLS forecasts under the most permissive cost grid. The dashed line is the taker fee the pre-registered design assumes.

The signal shows up almost every day

Per-day out-of-sample R² across 38 BTC and 37 ETH held-out days. Positive on 97% of them. Sign accuracy reaches 69 to 76% at 1 second on bars where the mid moved at all.

Per-day out-of-sample R² distribution, held-out days.

The founding hypothesis returned a null

The study exists because of a hunch that Shannon entropy of order flow carries information beyond the flow itself. It does not. Ablation ΔR² is roughly −0.0002 percentage points and day-level wins are never distinguishable from a coin flip.

The design was committed before any model was fit, so this null is reported with the same prominence as the positive results. A measurement study that cannot publish a null is not measuring.

Entropy ablation: ΔR² from adding entropy features. ΔR² of flow plus entropy versus flow alone, per symbol and horizon.
Validation and dataHow the pipeline was checked against ground truth, and what the dataset covers.Show

The pipeline was validated before it was trusted

Phase v1 scored every component against LOBSTER NASDAQ ground truth: zero reconciliation violations across 500,000 sampled book events, and the trade-sign classifier scored against true aggressor labels. Lee-Ready holds 90 to 94% accuracy under a 1 second quote lag. Classic stylized facts replicate.

Trade classifier accuracy by condition.
Accuracy by tick-size group and trade location.
Intraday volume profile, LOBSTER sample day. The classic U shape, replicated from the sample day as a pipeline sanity check.

44 months of tick data, pinned by hash

Dataset v2a

Binance USDT-perp quotes and trades, 2023-01 through 2026-08. 87 of 88 candidate day-symbols passed quality gates; every source file is recorded in a committed SHA-256 provenance manifest. The order-flow imbalance driving the study:

en  =  1{Pnb≥Pn−1b} qnb  −  1{Pnb≤Pn−1b} qn−1b−  1{Pna≤Pn−1a} qna  +  1{Pna≥Pn−1a} qn−1a\begin{aligned} e_{n} \;=\;& \mathbb{1}\{P^{b}_{n} \ge P^{b}_{n-1}\}\, q^{b}_{n} \;-\; \mathbb{1}\{P^{b}_{n} \le P^{b}_{n-1}\}\, q^{b}_{n-1} \\ &{} -\; \mathbb{1}\{P^{a}_{n} \le P^{a}_{n-1}\}\, q^{a}_{n} \;+\; \mathbb{1}\{P^{a}_{n} \ge P^{a}_{n-1}\}\, q^{a}_{n-1} \end{aligned}

Event-level order flow imbalance, Cont, Kukanov and Stoikov (2014), summed per bar. Implemented once in src/features/microstructure.py and reused across phases.

Trade-sign ACF surface, included day × lag. One row per included day. The persistence structure is stable across three and a half years.
Trade-sign autocorrelation, mean and 10-90 band.

Coverage, month by symbol

passed quality gatesexcluded: 60.2 s quote gap

87 of 88 candidate day-symbols pass. The one exclusion is the only gap in 44 months.

Four allocation rules, judged strictly out of sample

portfolio-optimization

Weights chosen on 2019-2022 daily data only, then held fixed and evaluated on 2023-2024 with quarterly rebalancing, gross and net of 10 bps per rebalance. Twelve large-cap US stocks; a fixed 60/40 SPY/AGG portfolio as the reference. The point is not the returns. It is the gap between what the optimizer predicted and what happened.

Growth of $10,000 over the evaluation window. Zero-cost curves shown; net-of-cost results differ by under 1% annualized and are reported in the repository.
Predicted vs realized volatility. Estimation error, the honest headline of portfolio optimization.
Rolling 60-day correlation vs SPY, asset × time.

Source repository, including the limitations the README states plainly: one regime, a universe picked with hindsight, and a flat cost model.

Mag 7 vs SPY through 2025, measured daily

mag7-quant-pipeline

An equal-weight basket of the seven mega-caps against SPY across 249 trading days of 2025: cumulative equity, 20-day rolling volatility, 60-day rolling correlation, and a daily regime tag. Descriptive analytics, not a backtest.

Cumulative equity, normalized to 1.0.
Rolling volatility and correlation.

Source repository

Systems and earlier work

AEGIS

Python, agentic trading platform

“The LLM proposes; deterministic code disposes.” An agentic trading platform where language-model output can never touch an order directly: every proposal must pass a deterministic policy gate before an executor sees it. Phases 0 and 1 of 10 are complete: a typed data layer over the Alpaca paper API with staleness-labeled quotes, option chains, and news, plus operator CLIs. Paper trading only, hardcoded. No live capital, and graded on operational correctness rather than P&L.

config.yaml+ envdata clientspaper=True, hardcodedTTL cachetyped modelsevery read stampedCLIscheck, snapshotLLM brainproposes onlypolicy gatedeterministic, disposesexecutorpaper brokerplanned (phases 2+):
  • 76 passing tests, fixtures only, no live API calls in the suite
  • Free-plan data quirks handled explicitly: IEX feed fallback, delayed options feed, staleness labels
  • Explicitly no live capital; the execution interface exists so the constraint is architectural, not aspirational

Repository

Kalman pairs trading

C++17, Eigen

A 2-state Kalman filter tracks the intercept and hedge ratio of a simulated cointegrated pair; the innovation z-score drives a mean-reversion state machine, with rolling OLS as the baseline. Built for the tick path: zero heap allocations (enforceable with Eigen’s runtime malloc assert), LDLT factor-and-solve instead of matrix inversion, Joseph-form covariance updates.

  • On a deterministic synthetic tape with a mid-run hedge-ratio break, the filter tracks the true ratio with 7.8x less error than rolling OLS
  • All numbers come from the simulator and the README says so; no market claims
  • Streaming risk metrics in O(1) memory: Welford Sharpe, running max drawdown

Repository

SMA backtest

Python, first project

The earliest project here, kept as a record of the starting point. One notebook: SPY 2015-2024, a 50-day moving average, long when price is above it, in cash otherwise, signals lagged one day. The strategy finished at 1.78x while buy-and-hold finished at 3.40x. No costs modeled, one parameter, in sample. It underperformed, and that is the result.

Repository

GARP policy letter

Writing, model risk, July 2026

A letter to the CEO of the Global Association of Risk Professionals proposing formal practitioner guidance on AI model diversity and vendor concentration disclosure, housed in GARP’s existing Risk and AI certificate. The argument: firm-level specialist agents do not address industry-level concentration, because firms building excellent narrow tools on the same few foundational models still fail together. Concentration in AI architecture should be treated like concentration in positions: measured, disclosed, managed.

About

Finance student at George Mason University working on market microstructure. I work in Python (pandas, NumPy), C++ (Eigen), SQL and Git.

github.com/yushingtoncity