The method

Three commitments apply to every equity strategy book. They are written down before the first trade, not after the result.

Looking for the system diagram? See the service flow on the Overview.

01 Pre-register — the rules, a decision date and a kill rule are written before the first trade. no moved goalposts
02 Trade forward — paper books trade live prices with modeled costs and no hindsight; every trade is time-stamped as it happens. backtests don’t count
03 Graduate or die — an equity book clears the graduation card below, or it is closed and stays on the public scoreboard as retired. §19 card

The graduation card (§19)

All six must hold at once for a paper equity book. The numbers are the constants the backend checks, not marketing copy. The BTC Trend paper book (42) is to be assessed for rule fidelity under a bounded protocol, not by this card (registry A27); that assessment has not started.

CriterionBarWhy it is there
Track record≥ 90 trading daysA short window cannot separate skill from luck
Sharpe ratio≥ 1.0Return per unit of risk, not raw return
Hit rate≥ 55% with ≥ 8 closed tradesA thin sample is never judged
Max drawdown< 15%Survivability
Regression alpha vs SPY> 0Return the benchmark does not explain; raw “beat SPY” is display-only
vs random portfoliosPaper return > 0 and above the 95th percentile of random-portfolio trialsMust beat luck drawn from the same universe

Clearing the card is not permission to trade real money. After it come, in order: a second, non-overlapping window of ≥ 90 trading days that replicates the result; execution-readiness and shadow-live qualification; and signed owner approval. Only then can a live book exist. As of 2026-09-30 no book has graduated and nothing trades live.

How the discipline works

The same four steps for every equity strategy book. Nothing runs indefinitely.

1 Backtest — every policy idea replays offline against historical decisions and bars before it touches a book. offline
2 Pre-register — a survivor gets a registry row: the question, a decision date, and the rule that kills or promotes it. registry
3 Paper soak — it trades paper against live prices, after modeled costs. paper only
4 Graduation read — on the date, the pre-committed criteria decide; a winner must replicate in a fresh window before real money. §19 + replication

Where the AI sits

The AI here doesn’t pick stocks. It labels the world so the math can. We tried the other way; it lost, and we published that (finding 04 below).

What agents do

  • Risk labeler → flags observable risk events; strategies that enable the risk-flag gate refuse buys on fresh high-severity flags
  • Retro narrator → nightly trade-level review, never on a request path
  • Macro narrator → writes the macro narrative; the regime itself is computed deterministically
  • Theme generator → search content, never a signal

What the quant does

  • Rank → LightGBM ensemble over the feature substrate
  • Gate → deterministic rules, pre-registered
  • Size & exit → rule-based sizing, rank-percentile and stop exits
  • Qualify → the §19 card, then a second window, before real money

Who grades whom

the gate
  • LLM reviewer on every decision → net-negative alpha, removed 2026-06-14
  • Crypto LLM verdicts → inverted, killed 2026-08-12
  • LLM picker → forward-graded daily against a quant control since 2026-09-11; grades exist, a verdict does not yet
  • Registry → a decision date and a kill rule for every A/B comparison

Generative AI never reaches an order without passing a pre-registered gate.

There is no LLM in the verdict path: verdicts are pure quant. Model names are vendor aliases and can be re-routed by the vendor.

What has been learned so far

Four findings from the research store and the paper books. None of them are flattering. Each is dated; none is a forecast.

01 · The classic factors lost to doing nothing clever

research store, summer 2026

Nothing long-only that was tested beat simply buying every eligible name equally on the long-history research store. The bar was never the S&P 500; it was the equal-weight eligible universe, and that bar has held so far.

02 · The model’s top decile trailed its own middle

2026-06-15 → early Aug 2026

Over a 39-trading-day window the top-ranked decile returned less than the middle-ranked names over the following 20 sessions, on 36 of 39 days. A five-hypothesis study ruled out sector, volatility, extension, single names and data artifacts: the score itself inverts inside the top half. The gap faded across the window, and overlapping returns leave roughly two independent observations — a well-checked description of one episode, not a statistically tested permanent property.

03 · A Bitcoin trend simulation showed lower drawdown, with limits

historical simulation · paper only

The rule holds Bitcoin above its 100-day moving average and cash otherwise. In a historical simulation from 2014-07-20 through 2026-09-29, maximum drawdown was 64.9% versus 83.6% for BTC-hold before costs. With modeled costs of 0.6% per trade side, strategy drawdown was 69.3% and annualized return was 43.1%, versus 49.5% for uncosted BTC-hold. Drawdown was not lower in every period. The simulation assumes fills at the same final close used to calculate the signal; it does not establish executable live prices. Book 42 is paper only, and its original execution soak failed because of a late entry. These historical results do not constitute forward validation.

04 · The AI stock-picker was fired. Twice.

2026-06-14 · 2026-08-12 · 2026-09-11

An LLM reviewer sat on every decision and produced net-negative alpha, so it was removed on 2026-06-14. The crypto LLM verdicts came out inverted — the “buy” bucket (646 calls) performed worst — and were killed on 2026-08-12. A self-audit on 2026-09-11 then found one of our own claims false: the picks were being logged, but nothing was scoring them. The scorer now exists and runs daily; it has produced grades, not a verdict on any picker.

Nothing has graduated to live trading yet. That’s the point of the gates.

What runs today

  • Paper books trade the engine’s decision row under pre-registered rules. As of 2026-09-30 the public standings list five active paper books (Pure Quant, Pure Quant · Trailing, Insider, Analyst Consensus, BTC Trend) and four retired. BTC Trend is to be assessed for rule fidelity under a bounded protocol that has not started, rather than for performance.
  • Nothing is live. No real orders are placed; every figure on the scoreboard is Paper.
  • The AI labels, narrates and flags risk — it never sets a verdict, though a high-severity risk flag can block a paper buy.
  • Every A/B comparison has a decision date and a kill rule, set before it starts. The BTC Trend book has a pre-registered fidelity check with a fixed session limit instead.
  • The public scoreboard at www.vibebullish.com/autopilot shows every paper book of the current epoch, including the retired ones. Engineering shadow books run beside the measured books and are excluded from §19.

Disclaimer

VibeBullish is an educational research project. The strategies shown are paper strategies: they trade simulated capital against real market data and pay modeled costs, but no real orders are placed. Nothing on this site is investment advice or a recommendation to buy or sell any security. No performance is claimed beyond what the paper scoreboard visibly shows, and past paper results say nothing about future results.

Verified 2026-09-30 against vibebullish-docs@449ec5f (experiment registry amendments A22 and A27; the Bitcoin simulation results in research/btc-trend-gate-2026-09-29 in the repository) and vibebullish-backend@a15ecf7 (graduation constants in book_equity_status.go; LLM roles in agent_roles.go), plus one anonymous read of the public standings at 2026-09-30 20:14 UTC (five active books, four retired). Runtime observations are dated inline; a source sha does not vouch for deployed configuration or database contents.