RESEARCH
how the market actually moves · the day 2 white paper
v1 · 2026-08-05
If you came here for a hack
There isn't one. That's the finding, not a disclaimer. We measured what actually predicts price movement on 25 years of real data across 12 world markets plus Thailand's SET, and what survived is a short list — eight edges, most of them slow, all of them printed with the sample size that should make you suspicious of them. What did not survive is a much longer list, printed too, because knowing what doesn't work is the actual product.
If "hack the market" means finding a lever nobody else has pulled: no. If it means replacing hope with arithmetic: yes, and here is all of it, in one place, for the first time.
The three honest questions
Everything in Day 2 answers one of three questions, and refuses to blur them together:
- Is this stock cheap relative to what it earns and owns? — Graham arithmetic. Answerable for any listed company with positive earnings and book value. THE LENS.
- Is stress about to travel from one market to another? — volatility spillover. Answerable, but it tells you about risk, never about direction. THE MAP.
- When have specific, named conditions historically preceded specific, named returns — and how many independent times did we actually see it? — the directed edges. Answerable for a short list, with wide error bars. THE MAP's edge cards, THE PLAN.
Conflating these three is how most retail finance content is built: a cheap stock chart with a trend line implies question 3 answered question 1's question. Day 2's whole design — three sizes of type, one accent color, a graveyard next to every live claim — exists to keep them separate.
What actually predicts, ranked
evidence.md, re-measured here
2.1 · the network methods: what each one is honestly allowed to say
| Method | Claims about | Status |
|---|---|---|
| Diebold-Yilmaz volatility spillover (VAR + generalized FEVD, Pesaran-Shin) | Contagion risk — not return direction | REPLICATED canonical in central-bank practice (Diebold & Yilmaz, financialconnectedness.org), re-measured live on our own 12-market panel — see §3 |
| Graphical Lasso / partial correlation | Portfolio-variance reduction, not returns | CITED Goto & Xu 2015 JFQA; Lee & Seregina 2023 arXiv:2011.00435 — not yet wired into Day 2 |
| Granger-causality, pairwise | Individually unreliable — common-driver contamination | REFUTED-BY-OUR-DATAin principle we refuse to draw pairwise Granger arrows on /map at all (only aggregate density is defensible, per Billio, Getmansky, Lo & Pelizzon 2012 JFE) |
| Transfer entropy | Descriptive lead-lag confirmation, fragile out-of-sample | CITEDnot built estimation burden too high for the data we hold |
| MST/PMFG correlation topology | Peripheral nodes diversify better than central ones, out-of-sample | CITED Pozzi, Di Matteo & Aste 2013, Scientific Reports — our own topology layer is roadmap, not yet MEASURED |
2.2 · the eight edges, all measured on our own lake
Sample: /Volumes/Data/marketdata, 2001-01-01 → 2026-07-24 (SET index bars end 2026-07-17 — see §8 on data honesty), computed by backend/analysis/influence.py, republished daily.
| # | Edge | Status | Number | n | Skepticism |
|---|---|---|---|---|---|
| 1 | S&P 500 (yesterday) → SET (today) | REPLICATED | β 0.24 · corr 0.25 · 58% sign-match | 5,457 | Explains the open — priced in by the open. Interpretable, not tradable. |
| 2 | VIX ≥ 30 (panic) → SET, forward 90 days | REPLICATED | +12.6% vs +3.0% baseline, 75% hit rate | 547 daily obs (~26 independent episodes) | Wide error bars from so few independent episodes; 2008 kept falling for months after the first spike. |
| 3 | WTI crude → SET energy complex (PTT/PTTEP/PTTGC) | MEASURED | same-day corr 0.18, overnight β 0.17 | 5,603 | Sector edge only — oil→SET index is sign-ambiguous (Thailand is a net oil importer with an energy-heavy index). |
| 4 | USD/THB (baht weakness) → SET | REFUTED-BY-OUR-DATAat daily frequency | same-day corr −0.08 | 5,147 | Weaker than its reputation. The foreign-flow mechanism is real but works over weeks/regimes, not days. |
| 5 | US 10y − 3m yield curve → recession odds → world equities | CITEDcontextualized | today's spread +0.87pp; inverted 13% of days since 2001 | 6,421 | Recession probability, NOT equity timing — 2019's inversion said nothing useful about 2019 returns. |
| 6 | SET vs its own 200-day average → forward 90 days | MEASUREDweak | above-MA +3.7% vs below-MA +2.8% | 6,131 | A regime filter, not alpha — overlapping windows overstate the apparent sample. |
| 7 | SET 12-month trend (skip last month) → next month | REFUTED-BY-OUR-DATA | up-year fwd +0.38% vs down-year fwd +1.28% | 5,948 | The replicated cross-sectional EM momentum finding (NBER w31839, 14/21 EM markets) does not survive at the index level here — cross-sectional stock-vs-stock ranking is a different, unbuilt measurement. |
| 8 | Calendar Nov–Apr → SET monthly return ("Halloween") | REPLICATED | +0.88%/mo vs +0.49%/mo May–Oct | 306 months | The one calendar effect with a genuine 323-year, 109-market out-of-sample test (Zhang & Jacobsen 2021) — no accepted mechanism, never a standalone trade. |
2.3 · the graveyard — measured dead, on purpose displayed as prominently as the live edges
| Claim | Status | Number |
|---|---|---|
| "Buy the dip when RSI < 30" | MEASUREDdead on SET | forward 90d +1.3% vs +3.3% just holding — edge −2.0pp, n=724 |
| Pairwise Granger-causality arrows between named assets | REFUTED-BY-OUR-DATAin principle | spurious under a common world-equity-factor driver — only aggregate density survives, as a stress thermometer |
| Post-earnings announcement drift, large caps | DEAD | extinct since ~2006 (Martineau, "Rest in Peace PEAD") |
| Pre-FOMC announcement drift | DEAD | disappeared after its own 2015 publication (Lucca & Moench 2015 JF; death documented 2020) |
| "Low VIX means sell" | REFUTED-BY-OUR-DATA | no predictive content at low readings — only the panic extreme carries signal |
Where shocks travel right now
MEASURED DY window ending 2026-07-16: total connectedness 42.7, the 19th percentile of its own trailing 3-year history — calm, shocks staying local rather than cascading.
net transmitters vs. receivers this window (positive = net exporter of volatility to the system)
KOSPI sits almost at the apex, tied with S&P — surprising, possibly a VAR artifact from session-overlap ordering or high-beta tech concentration, and printed as a caveat rather than quietly corrected, because the measured number outranks the story. Hong Kong and Singapore — the entrepôt aristocracy of Asian finance — are measured net receivers this window, not transmitters. Prestige rank and flow rank are not the same thing. That disagreement is not a bug; see §4.
The full web: 31 nodes, 90 links — 15 market indices, 11 macro instruments (gold, DXY, Brent, WTI, VIX, yields, four FX crosses), 5 SET sector baskets built from our own 25-ticker universe (banks, energy, telecom/tech, consumer, property). Two link types, never conflated: DY spillover + the 8 directed edges (strong ropes, causal-adjacent), and pairwise |correlation| ≥ 0.25 over the trailing 250 days (thin ropes, descriptive co-movement only — a diversification map, never causality). A node wrapped in ropes is crowded; a node with few is a diversifier.
The geography of money
G. William Skinner's central-place theory of rural Chinese marketing systems (Skinner 1964–65, Journal of Asian Studies 24; Skinner ed. 1977, The City in Late Imperial China, summarized in Daniel Little's 2008 memorial essay, The China Beat #271) describes an orderly hierarchy: a few central metropoles, intermediate market towns, and many standard markets at the periphery — nested, with goods and information flowing predominantly downward and mostly within physiographically bounded macroregions rather than across them.
applied honestly to §3's measured hierarchy (full synthesis in docs/research/skinner-geography.md)
- Central metropolis = highest net transmitters (S&P, KOSPI)
- Intermediate market towns = mid-band transmitters (FTSE, Nikkei)
- Standard markets / periphery = net receivers (NIFTY deepest; SET weakly coupled on both sides — a standard market oriented upward via the overnight edge, but off the main circuits)
- Marketing-area nesting = the DY spillover shares themselves (KOSPI⇄ Nikkei 15.2/13.7%, FTSE⇄DAX 11.1/10.2%)
- Macroregions trading mostly internally= an intra-East-Asia block and an intra-Europe block dominate the strongest pairs; the metropolis (S&P) is what discharges across blocks
- Two coexisting hierarchies, administrative vs. economic(Skinner's sharpest methodological point) = exactly the HK/Singapore finding above. Skinner built this distinction because a place's rank in the prestige hierarchy need not match its rank in the measured-flow hierarchy. We display the measured one, which is the whole discipline in one sentence.
Where the analogy strains, on the record: volatility is not goods (the hierarchy is of stress transmission, never return direction); Skinner's geography was quasi-stable over centuries, ours is a 250-day snapshot that reorganizes violently in crises; the hexagons do not survive (financial space has no transport-cost metric — a hexagonal UI would be decoration, and this project refuses decoration); the periphery's normative valence flips (extracted and poor in Skinner, prized in portfolio theory — Pozzi 2013's peripheral-diversifier finding, still CITED not yet MEASURED here).
What we chose not to build, and why
fincept-terminal.md, tradingagents.md
FinceptTerminal(29.2k★, AGPL-3.0, ideas studied not code copied) proved two things worth inheriting: a DataHub-style single pipe with per-topic freshness TTLs is the right answer to "twenty widgets share one rate-limited API," and a four-score valuation strip (Graham/Altman/Piotroski/Beneish, shown as auditable checklists) is the right shape for "is this stock cheap." It also proved a cautionary tale: 57 screens including maritime vessel tracking and RL trading labs, a cloud dependency for its own core valuation math that is currently brokencrippling the one screen that mattered, and 37 AI "persona" agents that are prompt wrappers over the same underlying data. Day 2 kept the auditable-checklist idea (THE LENS showcase) and refused everything else.
TradingAgents(94.7k★, Apache-2.0, actively maintained) is the more important refusal, because its headline claim is false. Its own paper reports a 62-trading-day backtest on 3 mega-cap tickers with Sharpe 5.60–8.21 and no transaction costs modeled. An independent replication (ACM AI & Fintech 2026, doi 10.1145/3800973.3801029) ran the same framework on GOOGL May–Jul 2025: GPT-4o scored 15.8%±4.2%, Qwen3 scored 18.1%±2.8%, buy-and-hold scored 19.1%— both configurations lost to doing nothing, and the run-to-run standard deviation (~9%) means roughly half of any single reported number is sampling noise. The broader multi-agent-debate literature agrees this isn't a fluke: five methods across nine benchmarks fail to consistently beat a single well-prompted model at equal compute (arXiv 2502.08788). DEADindependently confirmed
What Day 2 kept from it: bull-case/bear-case framing before a verdict, structured Pydantic-style output over free-form chat, and a verdict journal with outcome reflection — genuinely useful patterns, extracted from a mechanism whose central performance claim does not survive replication. What Day 2 refused: any autonomous multi-agent trading pipeline, and any continuous LLM analysis (an on-demand ~$1/click verdict is economically sound; continuous SET50 coverage at $1–3k/month for a signal that's ~50% noise is not).
The showcase, live
MEASURED daily by backend/publisher/publish.py::build_showcase(): Graham arithmetic (GN = √(22.5 × EPS × BVPS), MoS = (GN−price)/GN×100) computed identically to app/src/lib/graham-lens.ts, over our 25-ticker SET universe (lake closes) plus 8 curated global names (live-fetched).
Coverage today: 31 of 32 scored — one skipped, and the skip itself is a finding.Berkshire Hathaway (BRK-B) reports its book value in the wrong share-class units through the fetch path we use; a naive computation produced a fabricated +97.5% margin of safety before the unit-sanity gate caught it. It is now excluded with the reason on record — "provider book value in wrong share-class units — refusing to fabricate a Graham number" — rather than shipping a number that looks precise and is false.
The one decision
MEASURED from the lake: SET index price CAGR 7.29%/yearover 25.5 years (2001-01-03 → 2026-07-17, 272.03 → 1,639.04) and S&P 500 price CAGR 7.10%/yearover 25.6 years (1,283.27 → 7,411.98). Both figures are price-only — dividends excluded, and the extract says so explicitly, because SET's dividend yield has historically run near 3%/year and omitting it is the conservative choice, not an oversight. The US 10-year treasury yield, 4.68% as of 2026-07-24, is the honest bond-market reference point; there is no 25-year total-return bond series in the lake, so the PLAN screen says exactly that rather than filling the gap with a guess.
Against those two measured growth rates, the compounding arithmetic in app/src/lib/plan-math.ts (unit-tested, 9 passing cases, pure function, throws on nonsense inputs rather than silently returning garbage) answers one question honestly: at your current spending and savings rate, does money sitting in a Thai savings account (Bank of Thailand reference rate, cited, ~0.25–1.5%, against inflation of ~1–2%) outlive you — or does it run out first? For most inputs, the historical gap between "bank" and "measured market CAGR" spans decades. That gap, not a stock tip, is the one decision this whole system exists to surface.
Data honesty
- The SET index bars in our own lake end 2026-07-17— nine days stale as of this white paper's date. Yahoo's ^SET.BK feed lags or drops intermittently; the WORLD page's red staleness dot on the SET tile is this fact surfacing correctly, not a bug. Every measured figure above that uses ^SET.BK inherits this staleness honestly through its own windowEnd/asOf stamp.
- KOSPI's near-apex transmitter rank (§3) is flagged, not corrected, because correcting a measured number to match intuition is exactly the failure mode this whole project exists to refuse.
- VIX-panic's n=547 is 547 daily observations, not 547 independent panics— roughly 26 distinct episodes in 25 years. The card says so. Treat every "n=" in this document the same way: ask whether the observations are independent before trusting a percentage.
- The MST/topology core-periphery layer is CITED, not MEASURED— it is Phase-2 roadmap. Until it is computed from our own lake, the "periphery diversifies" claim rests on Pozzi et al. 2013 alone, not on Day 2's own data, and is labeled that way everywhere it appears.
What this white paper is for
Not to be read end to end by most people who use Day 2 — THE MAP, THE LENS, and THE PLAN exist so the measured findings above reach someone in fifteen seconds, not fifteen pages. This document exists so that every claim any of those three screens makes traces back to a number, a sample size, and a citation, in one place, checkable by anyone with the same public data and the same lake.
That traceability — not a clever signal — is the actual answer to "how to hack the market." There isn't a shortcut. There is only the discipline of knowing, precisely, which of your beliefs about markets are measured and which are just repeated.
docs/WHITEPAPER.md · docs/METHODOLOGY.md · docs/ARCHITECTURE.md — day2.nonarkara.org