Pi² Industrial Innovation Project · ESILV × Ginjer-AM

A Causal Approach to Bitcoin Performance Modeling

Applying López de Prado's Causal Factor Investing Framework to identify the genuine causal determinants of Bitcoin price behavior across investment horizons.

34Features Tested
10Causal Drivers Validated
0Macro Survivors
4Discovery Algorithms
01 — Context

Research Question

Bitcoin has attracted significant institutional interest, yet a fundamental question remains unresolved.

🔗 The Problem

Traditional correlation-based analysis is unreliable: correlations between Bitcoin and other assets are unstable over time and highly sensitive to the investment horizon and macro-financial context. Should Bitcoin be in a core portfolio for diversification, or confined to a satellite allocation?

🎯 Our Approach

We shift from associational analysis toward causal inference, applying Marcos López de Prado's 3-step framework — the same methodology used at ADIA Lab (Abu Dhabi's sovereign wealth fund). This lets us separate genuine causes from spurious correlations.

02 — Methodology

López de Prado's 3-Step Framework

Most finance papers stop at Step 1 (correlation). We go all the way to Step 3 (falsification).

STEP 1

Phenomenological

Observe statistical associations across 34 features and 4 pillars. ADF + KPSS stationarity tests. Correlation heatmap reveals unstable associations.

→
STEP 2

Theoretical

Propose causal structures using 4 complementary algorithms. Consensus rule: feature detected by ≥ 2/4 methods qualifies as causal candidate.

→
STEP 3

Falsification

Test causal claims via DoWhy: placebo treatment, random common cause, data subset refutation. Only robust features retained → sparse, validated DAG.

03 — Data Architecture

Four-Pillar Feature Organization

34 features from Bloomberg and CoinMetrics (2015–2025), organized into four thematic families.

On-Chain

9 features
CapMVRVHashRateMining Diff.Miner Rev.TxCntAdrActCntTX VolumeNVTSplyCur

Macro

12 features
DXYVIXMOVESPXCCMPRTYUS 3MGoldOilM2FARBAST2Y10

Sentiment

5+ features
Fear & GreedStablecoin MCapExchange Bal.Google "Bitcoin"Google "Coinbase"

Crypto Market

5 features
ETHBNBXRPLTCCFTC Futures
Preprocessing: log-returns for prices, first-differences for bounded ratios · forward-fill only (no look-ahead bias) · exclusion of BTC-derived features · weekly resampling for Google Trends compatibility
04 — Discovery Algorithms

Robustness via Multi-Method Consensus

Four independent algorithms with different assumptions, applied to the same weekly-resampled dataset.

PC Algorithm Constraint-based

Tests conditional independence via Fisher-Z on partial correlations. Removes edges that are conditionally independent. Contemporaneous structure only.

X ⊥⊥ Y | Z ⟹ remove edge
α=0.05 · 4 causes → BTC (ETH, BNB, CapMVRV, Miner Rev.)

NOTEARS Optimization

Continuous optimization: learns a weighted adjacency matrix W. Produces sparse DAG directly. Most aggressive filter — fewest edges retained.

min ‖X − XW‖² s.t. tr(eW∘W) − d = 0
λ₁=0.1 · 2 edges (BTC→ETH, BTC→CapMVRV)

PCMCI Temporal

Partial correlations at each lag, conditioning on all other lagged variables. Separates contemporaneous from lagged effects — captures time structure.

ρ(Xt−τ, Yt | Z) for τ ∈ {0, …, τmax}
τ_max=4w · 26 significant links (most detections)

Granger Causality Predictive

F-test on nested VAR models — does X's past improve Y's forecast? Bidirectional testing. Sensitive to lag selection.

F-test: VAR(Y|Ylag) vs VAR(Y|Ylag,Xlag)
max_lag=8w · Many bidirectional links
Consensus rule: ≥ 2 out of 4 methods → 10 candidates from 34 tested.
Each then undergoes DoWhy falsification (placebo, random common cause, data subset).
05 — Validated Results

Final Causal DAG

10 features survived ≥ 2/4 discovery methods + DoWhy refutation tests. Zero Macro variables survived.

Method Consensus Matrix

Feature Category PC NOTEARS PCMCI Granger Count DoWhy
CapMVRVOn-Chain✓✓✓—3✓
Miner RevenueOn-Chain✓—✓✓3✓
Mining DifficultyOn-Chain—✓✓—2✓
Est. TX VolumeOn-Chain——✓✓2✓
ETHCrypto✓✓✓—3✓
BNBCrypto✓—✓✓3✓
XRPCrypto——✓✓2✓
Stablecoin MCapSentiment——✓✓2✓
Google "Coinbase"Sentiment——✓✓2✓
Google "Bitcoin"Sentiment—✓✓—2✓
Eliminated: DXY, VIX, SPX — each detected by only 1/4 methods. Zero Macro variables survived.

Validated Causal Structure — 10 Drivers → BTC

ON-CHAIN CRYPTO MARKET SENTIMENT CapMVRV 3 methods Miner Revenue 3 methods Mining Difficulty 2 methods TX Volume 2 methods ETH 3 meth. BNB 3 meth. XRP Stablecoin MCap "Coinbase" search "Bitcoin" search BTC Returns Macro variables (DXY, VIX, SPX, Gold, Oil, M2…) tested but eliminated — detected by only 1/4 methods
Explore the Full Interactive Causal Graph (PC Algorithm)

Interactive Plotly graph — hover over nodes and edges to explore the complete causal structure discovered by the PC method.

06 — Multi-Horizon Analysis

No Universal Driver

Dominant factor families rotate across investment horizons — single-specification models are misleading.

Clustered MDA Importance by Family

Normalized positive Mean Decrease Accuracy (%) from Random Forest per pillar. Source: horizon_analysis.py

On-Chain
Macro
Sentiment
Crypto Mkt
T+1
49.2%
27.2%
23.0%
0.6%
T+7
15.7%
10.2%
64.6%
9.5%
T+30
26.2%
57.3%
1.8%
14.7%
On-Chain dominates at T+1 (49.2%) → Sentiment peaks at T+7 (64.6%) → Macro emerges at T+30 (57.3%)
No single pillar dominates universally. Driver rotation is the core finding.
07 — Predictive Validation

Causal vs. Full Feature Set

A Random Forest classifier (Bear / Neutral / Bull) demonstrates that causal features act as natural regularization.

Horizon All Variables (138) Best Causal (4–8) Dummy Baseline
1d 33.6% 33.8% 33.0%
7d 33.1% 41.7% 30.3%
14d 34.7% 42.2% 28.8%
30d 31.9% 35.4% 22.7%
60d 18.1% 50.2% 12.9%

📉 Why 138 features fail

With ~400 weekly observations, the all-features model suffers from the curse of dimensionality: it overfits noise rather than learning signal. At 60d, it degrades to 18.1% — worse than the 12.9% dummy baseline.

📈 Why 4 features win

López de Prado's pipeline acts as domain-knowledge-driven feature selection. The 4-feature causal model avoids overfitting, reaching 50.2% balanced accuracy at 60d. At 7d: 66% Bear recall, 54% Bull recall.

08 — Regime Analysis

Macro Regimes as Exposure Filters

Although no Macro variable survived the causal pipeline, macro conditions still condition BTC returns — they're context, not cause.

BTC Forward Returns by Macro Regime (VIX × DXY median split)

0.1%
0.5%
Risk-on
Weak $
0.3%
0.5%
Risk-on
Strong $
0.4%
0.7%
Risk-off
Weak $
0.2%
0.7%
Risk-off
Strong $
14d returns 30d returns
30d returns are consistently higher than 14d across all regimes. Risk-off / weak USD shows the highest returns — contradicting a simple pro-cyclical narrative. This regime framework is designed as an intermediate step: as institutional investors enter the crypto market, new causal links to macro variables may emerge.
09 — Decision Framework

Operationalizing Causal Findings

A committee-ready decision process separating regime context from causal drivers.

1
Choose Horizon
Driver families rotate across horizons (Fig. 2). Select target investment horizon first.
▼
2
Apply Macro Regime Filter
VIX × DXY classification → context, not cause. Conditions the environment.
▼
3
Monitor Causal Drivers
On-chain + crypto market signals from the validated DAG. Transparent, explainable.
Output: Transparent, explainable exposure rules compatible with institutional governance.
10 — Conclusions

Key Findings

▸
No universal driver: Dominant factors rotate across horizons — single-specification models are misleading. On-Chain at T+1, Sentiment at T+7, Macro at T+30.
▸
Macro variables fail falsification: DXY, VIX, SPX each detected by only 1/4 methods. No Macro feature survives the full pipeline. This likely reflects the dominance of retail investors who don't trade BTC based on macro aggregates.
▸
On-chain + crypto market dominate: 7/10 validated drivers are blockchain-native (CapMVRV, mining metrics, altcoins). Bitcoin's price is driven by its own ecosystem.
▸
Altcoins are causal, not just correlated: ETH, BNB, XRP survive all 3+ methods and DoWhy refutation — their relationship with BTC passes causal falsification.
▸
Causal > full feature set: 4–8 causal variables outperform 138 features at all horizons beyond 1d. At 60d: 50.2% vs 18.1% balanced accuracy. Natural regularization through domain knowledge.
11 — Perspectives

Future Directions

🏛️ Institutional Adoption

The current causal structure reflects a retail-dominated market. As institutional players enter (ETFs, sovereign funds), new causal links to macro-financial variables — particularly liquidity proxies (Fed balance sheet, M2) — are likely to emerge, reshaping the DAG.

🔬 Non-Linear Methods

All 4 discovery algorithms assume linearity (Fisher-Z, linear SEM). Threshold effects and regime switches may require non-linear extensions such as kernel-based PCMCI.

⏱️ Temporal Resolution

Weekly resampling was necessary for Google Trends compatibility but loses intra-week dynamics. Daily-frequency causal discovery could reveal short-lived trading signals.

🧪 Out-of-Sample Testing

Validate the regime-aware framework on post-sample data and across other crypto-assets (ETH, SOL). Explore alternative regime definitions (growth × inflation à la Dalio's All Weather).