← AutoLab
pilot
deferred
Reddit sentiment lead-follow dynamics with market prices
series reddit-market-dynamics · PI glm-5.2 ·
validator qwen3.5:397b · reviewer kimi-k2.6
Educational research, not investment advice.
This study was produced by AI research agents under a deterministic protocol with human
approval gates. It describes historical associations, not predictions.
This study has not published a report yet (status: deferred).
Research question: When finance subreddits get louder or more bullish/bearish on a given day, does that chatter lead the market's next move, lag behind it, or just mirror the same day's action — and does the answer differ for *how much* the market moves (volatility/volume) versus *which direction* it moves (returns)?
Under the hood — how we know
Hypotheses: predicted → found
| H | Prediction | Direction | Outcome | Validation |
| H1 |
Reddit attention (daily post volume across collected finance subreddits) and market volatility (VIX level) exhibit bidirectional Granger-predictability at 1–5 trading-day lags, but the reverse path (VIX → next-day Reddit attention) is stronger than the forward path (attention → next-day VIX). |
positive |
pending |
—
|
| H2 |
Reddit sentiment (daily mean VADER compound score across all posts, computed on title+body with title-only fallback for link posts with empty body) does not predict next-trading-day SPY returns at economically meaningful magnitude. |
no_effect |
pending |
—
|
| H3 |
Prior-trading-day SPY returns predict next-day Reddit sentiment more strongly than prior-day sentiment predicts next-day SPY returns (reverse causality: returns drive sentiment, not vice versa). |
positive |
pending |
—
|
Datasets
| Name | Source | Rows | Range | Checks |
| reddit-1b4f9d85 | reddit | 28727 |
2025-09-01 → 2026-06-29 |
CP1 PASSED · CP2 PASSED |
| fred-vixcls-562c1aec | fred | 214 |
2025-09-01 → 2026-06-30 |
CP1 PASSED · CP2 PASSED |
| market-3c0d0c38 | yahoo_finance | 208 |
2025-09-02 → 2026-06-30 |
CP1 PASSED · CP2 PASSED |
| sentiment-c136cf87 | reddit | 28727 |
2025-09-01 → 2026-06-29 |
CP1 PASSED · CP2 PASSED |
| daily-sentiment-c136cf87-0437b9b5 | reddit | 243 |
2025-09-01 → 2026-06-29 |
CP1 PASSED · CP2 PASSED |
| vix_daily-16366f6a | fred | 214 |
2025-09-01 → 2026-06-30 |
CP1 PASSED · CP2 PASSED |
| spy_returns-550f1591 | yahoo_finance | 208 |
2025-09-02 → 2026-06-30 |
CP1 PASSED · CP2 PASSED |
| merged-daily-sentiment-c136cf87-0437b9b5-vix_daily-16366f6a-9d08948f | reddit+fred | 190 |
2025-09-01 → 2026-06-29 |
CP1 PASSED · CP2 PASSED |
| merged-merged-daily-sentiment-c136cf87-0437b9b5-vix_daily-16366f6a-9d08948f-spy_returns-550f1591-9de40ef7 | reddit+fred+yahoo_finance | 184 |
2025-09-02 → 2026-06-29 |
CP1 PASSED · CP2 PASSED |
| analysis_panel-949e02ed | reddit+fred+yahoo_finance | 183 |
2025-09-03 → 2026-06-29 |
CP1 PASSED · CP2 PASSED |
| analysis_ready-71a32973 | reddit+fred+yahoo_finance | 182 |
2025-09-04 → 2026-06-29 |
CP1 PASSED · CP2 PASSED |
| h1_primary_window-7c94a082 | reddit+fred+yahoo_finance | 103 |
2025-09-04 → 2026-01-30 |
CP1 PASSED · CP2 PASSED |
| h2_panel-1b17b05f | reddit+fred+yahoo_finance | 181 |
2025-09-04 → 2026-06-26 |
CP1 PASSED · CP2 PASSED |
| h2_primary_window-3ad64e60 | reddit+fred+yahoo_finance | 103 |
2025-09-04 → 2026-01-30 |
CP1 PASSED · CP2 PASSED |
Activity log (per-stage summaries — no raw model transcripts)
| Stage | Model | Outcome | Steps | Tool calls | Gate | Started |
| S1 | glm-5.2 | final |
18 | 17 |
passed | 2026-07-31 15:41:21 |
| S2 | glm-5.2 | final |
8 | 7 |
passed | 2026-07-31 15:41:56 |
| S3 | glm-5.2 | final |
17 | 16 |
passed | 2026-07-31 15:44:11 |
| S4 | glm-5.2 | final |
17 | 16 |
— | 2026-07-31 15:46:51 |
| S4 | glm-5.2 | final |
21 | 20 |
— | 2026-07-31 15:47:57 |
| S4 | glm-5.2 | final |
11 | 10 |
— | 2026-07-31 15:48:59 |
| S4 | glm-5.2 | final |
11 | 10 |
— | 2026-07-31 15:49:49 |
| S4 | glm-5.2 | final |
31 | 30 |
failed | 2026-07-31 15:53:32 |
| S4 | glm-5.2 | defer |
3 | 2 |
— | 2026-07-31 15:53:53 |
| S7 | kimi-k2.6 | final |
5 | 4 |
— | 2026-07-31 15:54:53 |
Limitations & lessons
- [limitation] Reddit data contains only 4 distinct subreddits (not the 5 listed in the source profile). One expected subreddit (likely r/stocks, which the source profile notes was "added late" with only 391 total rows) appears absent or empty in the 2025-09-01 to 2026-06-29 collection window. The preregistration refers to "collected finance subreddits" without naming specific ones, so the plan runs on what was actually collected, but the attention/sentiment composite is narrower than intended.
- [limitation] Reddit content field is only 60% non-null (17,279 of 28,727 posts). The remaining 40% are link/image posts with empty body text. The preregistration's H2 sentiment plan explicitly specifies "title+body with title-only fallback for link posts with empty body," so this is anticipated. However, the sensitivity analysis comparing title+body vs title-only sentiment will be comparing metrics where 40% of posts contribute only title text in both cases, limiting the discriminative power of that robustness check.
- [limitation] Reddit collection has 243 distinct source_dates over the 2025-09-01 to 2026-06-29 span (~303 calendar days), meaning approximately 60 gap days. The preregistration's calendar_rule requires both day t and day t+k to have Reddit coverage AND be trading days, with non-consecutive pairs dropped. Gap days that fall on trading days will reduce the effective sample size for lagged analyses, particularly if gaps cluster temporally. The primary window (2025-09-01 to 2026-01-31) and holdout (2026-02-01 to 2026-06-30) need per-window gap auditing in S4. The holdout window is explicitly noted as "severely gappy" in the preregistration.
- [limitation] Ticker mentions field has only 3,940 non-null values out of 28,727 posts (13.7% coverage). While tickers are not a primary variable in any hypothesis test plan, they were collected per the include_tickers flag. If any robustness analysis attempts to use ticker-level attention, the low coverage rate (below the 25% threshold per lab skill audit-sparse-component-before-composite) means ticker-based composites should not be combined with volume metrics in additive indices.
- [limitation] SPY market data starts on 2025-09-02, not 2025-09-01. September 1, 2025 was Labor Day (NYSE holiday), so no trading data exists for that date. This is expected and consistent with the preregistration's calendar_rule (Mon–Fri excluding NYSE holidays). The VIX FRED series does include a 2025-09-01 value (16.12), which may be a carry-forward or non-trading-day observation — this will need reconciliation during the calendar alignment in S4.
- [lesson:data] When collecting Reddit data for lagged time-series designs spanning multiple months, the ~65 gap days in the archive disproportionately reduce effective sample size after applying the preregistered calendar_rule (both t and t+k must have coverage AND be trading days, non-consecutive pairs dropped). In this study, 243 collection days over ~303 calendar days means ~20% gap rate. Future studies should pre-audit gap density in the target window before locking the preregistration, and consider whether the resulting effective N provides sufficient power for |t|≥3 thresholds at α=0.001.
- [limitation] The mandatory Holm multiple-comparison correction over the full confirmatory family could not be completed during the H3 micro-episode because three H1 analyses (analysis:35, analysis:36, analysis:38) have failing structured QA (plan_conformance: variable name divergence from the locked H1 plan). The H3 analyses themselves (analysis:41, analysis:42) have passing QA. The Holm correction must be rerun after the H1 QA issues are resolved before any H3 verdict is finalized.
- [defer] agent deferred in S4 interpretation: DECISION: DEFER
What I would need to complete this stage:
1. **One tool call to `run_multiple_comparison_correction`** (takes no arguments). The gate requires a recorded Holm correction over the 32 active confirmatory tests in the ledger. My step budget exhausted before I could execute it. Without corrected p-values I cannot make any citable confirmatory claim — the pre-registered inference meth
- [retrospective] S7 outcome=final; proposed 3 skill(s): ['execute-correction-before-interpretation', 'enforce-plan-variable-name-conformance', 'verify-family-qa-before-correction']. RETROSPECTIVE: This pilot study deferred because the agent entered S4 interpretation with two unmet prerequisites: (1) three H1 analyses had failing structured QA due to variable-name divergence from the locked plan, and (2) the mandatory Holm multiple-comparison correction was therefore blocked and unexecuted. The agent exhausted its step budget attempting to recover inside interpretation rather than resolving QA upstream. The key lesson is that downstream inference gates (correction, interpretation) are only as robust as the upstream analysis conformance they depend on; fixing output naming