zms618/research-idea-discovery
v3.5.0
Paper-grounded research discovery from failures, successful mechanisms, apparent contradictions, and real-world needs, with motivation and evidence checks.
v3.5.0 — Paper-grounded research discovery
- Add critical paper reading as part of scientific problem creation, with source-to-question evidence lineage and forward/backward related-work checking.
- Add four evidence-first generators: failure replication, reverse-engineering successful mechanisms, fair reconciliation of contradictory findings, deployment-to-science abstraction.
- Keep method-driven techniques as search heuristics only; any suggested problem still requires M0 and downstream novelty/necessity checks.
- Introduce scoped, optional PAPER_ANCHORED workflow gate with source-record replay; retain OPEN_PROBLEM and one-question paths.
- Add separate scientific-upside vs reproducible-speed assessments, concrete asset-check policy, conditional two-week GO/STOP decision plan, anti-fabrication tests and researcher desk simulation.
3.4.0 — Motivation portfolio, shortlisted gaps and CVPR aspiration
- Add high-impact portfolio-first discovery: 5–8 distinct motivations when evidence permits; evidence-/impact-/falsifiability-aware comparison and 2–3 justified shortlist entries.
- Add targeted gap investigation for all shortlisted motivations before selecting a final one; maintain backup motives and record negative results.
- Add optional CVPR Best Paper aspirational quality profile without treating awards as an official scoring rubric or changing other research domain settings.
- Add reproducibility-conscious structural portfolio/gap audit tool and regression cases; preserve existing M0/M1/M2/idea/method/experiment gates.
v3.3.0 — Motivation-First Gate (2026-10-09)
- Enforce M0 problem-first evidence review before idea seeds/data search/method synthesis; retain M1 and M2 later.
- Add
research-problem-scout, motivation-first protocol and structural CLI gate with guarded progression. - Fail-fast weak/problem-unverified motivation; provide neutral evidence probe or reframing, not a proposed algorithm.
- Add multimodal DG evidence-first researcher fixtures, legacy-state protection and regression tests.
- Preserve existing v3.2 method synthesis and all later evidence gates.
Changelog
3.2.0 — 2026-10-09
- Add research-method-designer with concrete operator/loss, TTA stepwise state updates and implementer-facing artifact requirements.
- Add two-role researcher/reviewer desktop simulation; force adversarial corrections to method and independent potential contributions rather than simply three named modules.
- Expose identifiability, oracle-label leakage, source-data constraints, missingness-vs-corruption, checkpoint mapping, source teacher failure and O(C^2) risks.
- Add realistic multimodal TTA worked example with unverified data access labeled, repository risks and E0/no-go conditions.
- Strengthen advisor-report checks and add researcher-utility self-audit with synthetic adversarial cases. Structural tests never prove novelty or reproducibility.
- Retain prior plugins' identity, integration defaults and original evidence/workflow features.
Changelog
3.1.0 — 2026-10-09
- Introduce plain-language Idea Explanation and Research Execution Planner as actual user-facing requirements.
- Force concrete dataset/backbone/checkpoint/baseline, falsifying E0, first-week actions, kill conditions or explicit verification blockers.
- Align research story, hypothesis and potential contributions with evidence status; no forced third contribution.
- Integrate supervisor report into MAIN skill (not only optional advisor skill).
- Add structure checks and regression tests; retain all v2.7 evidence gates and restore original three default prompts in the extension manifest.
Research Idea Discovery release history
v2.7.0 — Execution-evidence gates and regression checks
- P0: Separate dataset catalog/field declarations from evidence of actually reading a bounded local sample. Record SHA256 and parsed fields. Reject missing sample, HTML landing pages, placeholder URLs, and missing labels. Do not claim source/license/whole-dataset reproducibility from a sample check.
- P0: Guard REVIEW → PILOT with a recorded review decision tied to idea_id/revision. Guard PILOT → UPDATE with replayed run-log checks and measurements or documented failure. Reject dry runs, no-metric successes, and forged readiness flags.
- P1: Structured four-axis prior-art query coverage and source-locator records; no claim of exhaustive novelty. Coverage record needed before mechanism-logic phase when running local workflow.
- P1: Compare proposed and simple-baseline outcomes when evidence exists. Equal/worse results mark
requires_reauditand force renewed mechanism/motivation/story scrutiny. - P2: Local append-only idea ledger with duplicate signatures; revising source evidence invalidates dependent assessments. Adds reproducible adversarial fixtures.
- 66 tests pass locally. These check control logic, not academic originality or real-world availability without network/tool permissions.
v2.6.0
- Research Idea Discovery plus targeted manual ChatGPT Deep Research handoff and report return.