Legacy Deprecation Spec — ai_target_queue · cleaned_data · run_pipeline

2026-08-08. Source: full cross-repo reader/writer inventory (backend + pipeline). Status: EXECUTED 2026-08-17 (all phases) — kept as the record of what was removed and why.

Executive finding

All three subsystems are already ~90% dismantled: run_pipeline() was removed 2026-06-03 (tombstone at pipeline.py:3504), ai_target_queue has ZERO live code references in either repo (only a self-documenting no-op stub + stale docs/worktrees), and cleaned_data has zero writers — its remaining readers are dead paths returning empty (and one 6-hour scheduled job, run_ml, has thrown a swallowed AttributeError on every fire since the method it calls was deleted). No foreign keys anywhere. What remains is deletion, not migration.

Load-bearing — DO NOT TOUCH

market_news_loop (5-min, → market_catalysts) — CORRECTED 2026-09-11: this was never load-bearing. The note assumed a backend consumer from the loop's own "consumed by the backend" docstring; a cross-repo grep found the only readers of market_catalysts are a TTL deleter (expiring_catalysts_service.go:68) and a dashboard display listing (llm_usage_service.go:211-229) — no decision path, and MacroEnvironmentService never touches it (audit finding G10). The loop + classifier were deleted in pipeline PR #71; the table is retained pending a separate ops drop. Still load-bearing: sec_forward_loop + news_forward_loop (→ ticker_features_daily), refresh_stock_prices (10-min, sole writer of stock_prices — backend portfolio valuation reads it), _run_backtester_tasks + daily sweep + Slack summary, the inference boot guard, and pipeline_scheduler's remaining shell.

Phase 1 — zero-risk deletions ✅ EXECUTED 2026-08-17 (pipeline PR #46)

Stub + broken run_ml job + minimal-server /ml-analysis deleted (269 tests green). ai_target_queue turned out to already be absent from the DB — no drop needed. The macro-environment worktree pruned, along with 20 other stale merged worktrees and 57 dead local branches.

  1. Delete push_lightgbm_picks_to_queue() stub (pipeline.py:3413-3437).
  2. DROP TABLE ai_target_queue (unmanaged artifact; nothing recreates it).
  3. Prune the stale macro-environment backend worktree (holds the only surviving ai_target_queue service/handler files).
  4. Delete the broken run_ml 6-hour job (main.py:269-279,287) + run_ml_analysis (pipeline.py:280-296).

Phase 2 — dead API surface ✅ EXECUTED 2026-08-17 (pipeline #47 + backend #133)

Went further than spec: the ENTIRE PipelineDataService and TickerScoringService were caller-less — both files deleted (not just the six methods). 7 pipeline endpoints deleted; get_recent_sentiment_data helper retained (3 live FinBERT callers — Phase 3 scope). Dashboard + web pre-checked: zero references. 5. Backend: delete pipeline_data_service.go methods GetMLAnalysis / Get*StockAnalysis / GetStockComparison / AnalyzeSentimentBatch / GetDataSummary / GetCompletePipeline + the always-zero CleanedDataByType/TotalCleanedRecords fields; simplify calculateRealNewsScore (already returns hardcoded 0.5, discarding the call) and drop the pipelineService plumbing from market_data/ticker_scoring services. 6. Pipeline: delete endpoints /ml-analysis, /analysis/stock/*, /analysis/stocks, /analysis/compare, /analyze-sentiment, /data/summary, /data/quality. ⚠️ Before deploy: confirm vibebullish-data-pipeline-dashboard doesn't fetch /data/summary or /ml-analysis (only unchecked consumer).

Phase 3 — cleaned_data + raw_data storage ✅ EXECUTED 2026-08-17 (pipeline #48 + drop)

Models deleted before drop (per the create_all trap); DROP ran post-deploy — 563 MB reclaimed. Pydantic CleanedData DTO untouched (verified different class). DISCOVERY while verifying: DB is 23 GB — top offenders scanner_activity 10 GB, action_decisions 3.8 GB, lgbm_test_predictions 3.2 GB, ticker_data_snapshots 3 GB. Retention policies for these are the natural Phase 6 (not in original spec). 7. Delete src/individual_analysis.py + call sites; delete database.py readers (get_recent_data, get_all_recent_data, store_raw_data, cleanup_old_data, get_data_summary). 8. Delete the SQLAlchemy models BEFORE dropping tables — Base.metadata.create_all (database.py:230) silently recreates them on every boot otherwise. Then DROP TABLE cleaned_data, raw_data. Note: retention was never scheduled, so these tables are frozen at their 2026-06-03 size — reclaimable space. 9. CleanedData in data_sources.py:707-760 is a live in-memory DTO on the stock_prices path — replace the DTO, don't delete the model blindly.

Phase 4 — legacy mood/sentiment ✅ EXECUTED 2026-08-17 (pipeline #50)

GPTInsightsGenerator + MoodScoreCalculator + MoodScore deleted. Kept: KeywordSentimentAnalyzer + slimmed MLPipeline — /ml/sentiment/finbert has a live backend consumer (sentiment ensemble → AI rating) the inventory missed. 10. Delete KeywordSentimentAnalyzer / GPTInsightsGenerator / MoodScoreCalculator (ml_pipeline.py:29-281) once /analyze-sentiment is gone.

Phase 5 — hygiene ✅ EXECUTED 2026-08-17 (backend #135 + docs banners)

crypto_reddit_momentum stub marked fully dormant (both ends of its plan are gone); index/architecture/data-sources.html bannered as superseded. 11. Delete/annotate the crypto_reddit_momentum.go comment-stub; mark the six stale design docs superseded.

Generated from legacy-deprecation-spec.md · source last changed 2026-09-29 · regenerate with npm run build:docs