galbaz1/video-research
Video analysis, document extraction, cited research and knowledge workflows backed by the video-research MCP server.
Changelog
This file records notable changes by release. Versioned entries describe the behavior and validation reported at that time; use current guides for installation and provider configuration.
The format is based on Keep a Changelog, and versions follow Semantic Versioning.
Unreleased
[0.8.0-rc.1] - 2026-10-02
Added
- Native Codex plugin packaging with 22 bundled skills and an exact core MCP runtime pin. Codex marketplace installation supports the packed local candidate and the matching npm version without running the Claude installer.
- The multimodal programme's current core exposes 90 tools, including local image inspection/cropping, structured audio/video workflows, evidence-aware research, and optional adapters. Adapter prerequisites and evidence limitations remain documented; packaging does not qualify an external runtime or model.
Fixed
-
Corrected the packaged video-to-skill helper paths for native Codex installs while preserving the existing Claude layout.
-
Built-wheel verification now compares every MCP tool contract with candidate source instead of requiring an obsolete tool count, and retains configuration redaction and actual local image-crop checks.
-
Reconcile FFmpeg 6.1 MP3 encoder padding against exact packet timing so complete tracks export successfully while truncated tracks remain rejected.
-
Exclude the uncleared external renderer subtree from the core source archive and reject its payload in archive checks, including equivalent path spellings.
Acceptance boundary
- This is a core/plugin release candidate. Companion packages retain their own versions. Human source audit, held-out comparison, live-provider quality and blocked native-runtime qualifications remain open; no programme-wide acceptance or superiority claim is implied.
0.7.1 - 2026-09-30
Changed
- Rewrote the README and reader-facing documentation around installation, daily use, architecture, contribution, and release tasks. Added a documentation index and preserved historical findings, release records, and reporting commitments.
- Corrected configuration, knowledge-store, cache, installer, and companion setup explanations against the implementation. Companion package patch versions now include their rewritten READMEs.
- Removed model-specific claims from package descriptions and server discovery instructions; configured model selection remains the runtime authority.
- Repaired the upstream renderer submodule's public clone URL and pinned revision. Aligned companion refinement, rendering, and TTS selectors with that CLI. The core PyPI 0.7.1 archives were published before this source-only and companion follow-up; their immutable files remain unchanged.
Security
- Redacted the Semantic Scholar API key from read-only
infra_configureresponses, alongside the existing provider secrets. Expanded the secret-redaction regression check and built-wheel smoke to cover this field.
0.7.0 - 2026-09-29
Changed
- Default research and summary models now use stable Gemini 3.8 Flash. Unsupported minimal thinking requests return an actionable validation error. Explicit-cache workflows retain Google's supported generateContent endpoint.
- Upgraded all three Python packages to current stable dependencies, including FastMCP 4 and Google GenAI SDK 2, with explicit API-major bounds and refreshed locks. Expanded locked CI to companion packages and Python 3.14.
- Global installer registration now uses Claude Code user scope in
~/.claude.json, preserves custom environment/settings and local companion servers, and refreshes the server package during resolution. Updated Playwright MCP and moved to maintained Node runtimes. - Audited and simplified skill/provider guides, aligned onboarding with installed servers, inherited configured orchestration models, and removed obsolete media API instructions. Added a bounded plugin-maintenance skill.
- GitHub release creation now requires passing CI, synchronized versions, matching release notes and valid build artifacts. Tags outside main publish as prereleases. GitHub source publication does not upload to PyPI or npm.
Fixed
- Preserved SDK thought signatures and full content through video-session history and SQLite serialization.
- Read current Deep Research interaction steps, citations, and error lists; preserved typed failures instead of treating absent legacy outputs as success.
- Rejected malformed or escaping installer manifests, preserved unowned/customized files, and retained ownership evidence for modified obsolete files.
- Companion scene generation now isolates child settings, enforces terminal SDK success and budgets, and uses a valid configurable Claude model.
- Companion render acceptance requires a fresh nonempty artifact; cancellation stops and joins subprocess work. Injection filenames cannot escape projects.
- Removed committed documentation conflict markers and unsafe visualization cleanup of user directories or unrelated local processes.
0.6.1 - 2026-05-21
Changed
- Default Gemini model:
gemini-3.5-flashwiththinking_level="medium"(previouslygemini-3.1-pro-preview/high). Bothdefault_modelandflash_modelpoint at 3.5 Flash; preset opt-ins (best,stable,budget) remain available for Pro 3.1, Pro 3, and the legacy 3-flash-preview. The release rationale cited Google's benchmarks as showing better results than 3.1 Pro at about 4× speed and less than half the cost. .claude-plugin/plugin.jsonversion sync — was 12 versions stale at 0.3.3; now part of the documented version-sync policy indocs/PUBLISHING.mdalongsidepyproject.tomlandpackage.json.
Added
- Ollama vectorizer for Weaviate (
text2vec-ollama) — local embedding alternative to OpenAI / built-in (#61, community contribution). .github/workflows/release.yml— auto-creates GitHub Release with CHANGELOG section on anyv*.*.*tag push. Fills the gap that left v0.4.0 through v0.6.0 without release pages (now backfilled).
Security
- 31 Dependabot alerts resolved (3 critical, 10 high, 14 medium, 4 low) via
uv lock --upgrade. Includes fastmcp 3.0.2 → 3.3.1 (Gemini-CLI command injection), python-multipart 0.0.22 → 0.0.29 (DoS), python-dotenv 1.2.1 → 1.2.2 (symlink following), cryptography 46.0.5 → 48.0.0 (buffer overflow), authlib 1.6.8 → 1.7.2 (CSRF + OIDC). google-genai bumped 1.65 → 2.5.0 (major; API remained compatible with this project's usage, with 781 tests reported passing).
0.6.0 - 2026-03-25
Added
- Media production skills — 5 new plugin skills for end-to-end AI video production:
tts-production— ElevenLabs TTS voice-over with API patterns, voice presets, and FFmpeg audio recipesffmpeg-production— Video/audio processing reference with post-processing chain, platform presets, and codec selectionvideo-generation— Provider-agnostic video generation (Veo/Sora) with selection matrix and draft-to-final workflowvideo-production— Cinematic multi-shot orchestration with style anchors, 4 chaining patterns, and frame-level QAimage-generation— Style anchor and prompt optimization for mcp-image (Subject-Context-Style structure)
- Skill descriptions stayed under 200 characters with negative qualifiers; bodies stayed under 2,000 words, with detailed material in
references/(Level 3 progressive disclosure).
Changed
- Plugin skill count: 7 → 12
- FILE_MAP entries: 34 → 42
0.5.0 - 2026-03-17
Added
- Semantic Scholar integration — 5 new tools:
research_paper_search,research_paper_details,research_paper_citations,research_paper_recommendations,research_author_search - AcademicPapers Weaviate collection with deterministic UUIDs
- Automatic knowledge graph extraction —
content_analyze,video_analyze,research_deep,research_web,research_document,content_batch_analyzenow auto-extract concepts and relationships - ConceptKnowledge + RelationshipEdges collections populated automatically
- S2-specific error categories (
S2_RATE_LIMITED,S2_NOT_FOUND)
Changed
- Collection count: 12 → 13 (AcademicPapers)
- Tool count: 28 → 33
Fixed
- Weaviate vectorizer config mismatch (
text2vec-weaviate→text2vec-openai)
0.4.4 - 2026-03-13
Fixed
- Vectorizer auto-detection — default without
OPENAI_API_KEYis nowweaviate(built-in embeddings) instead ofopenai, which had silently failed for Docker users without an OpenAI key
Added
- Startup warning — logs a warning when
WEAVIATE_VECTORIZER=openaiis active butOPENAI_API_KEYis not set - Env template keys —
WEAVIATE_VECTORIZERandWEAVIATE_AUTO_MIGRATEadded to the installer env template with documentation comments
Changed
- Weaviate setup skill — Step 3 now shows deployment-specific env blocks (Docker vs Cloud) with vectorizer guidance
- Knowledge Store docs — replaced broken
docker runcommand with docker-compose snippet matching the setup skill
0.4.3 - 2026-03-09
Fixed
_extract_reportturn.text fallback — Deep Research reports delivered viaturn.text(instead ofturn.content[].text) were silently lost; now both formats are captured- Transient 403 retry in
research_web_status— polling now retries up to 3 times with backoff on transient 403 errors rather than failing immediately - Concurrency guard on
research_web— prevents launching a second Deep Research task while one is active (API allows only 1 concurrent task per key); returns actionable error with the active interaction ID - Timeout heuristic in
/gr:research-deep— skill now warns after 20 min and suggests cancel+retry after 30 min of polling without completion
0.4.2 - 2026-03-09
Fixed
/gr:doctorSerena tool leakage — added explicit tool discipline section preventing the command from selecting Serena'sRead FileMCP tool instead of Claude Code's built-inReadwhen both are available in the session- Banned
model: haiku— replaced withmodel: sonnetin 4 commands (doctor,getting-started,models,explain-status)
0.4.1 - 2026-03-09
Fixed
- MCP server startup hang — added 2-second TCP reachability probe before MLflow setup. Prevents indefinite hang when
MLFLOW_TRACKING_URIpoints to an unreachable server (e.g., stopped MLflow instance) __init__.__version__— synced withpyproject.toml(was stuck on0.3.9)
0.4.0 - 2026-03-07
Added
/gr:advisorcommand — workflow advisor that recommends the optimal/grcommand for any task. Checks prior work via knowledge store before recommending. Three invocation channels: explicit command, auto-invoked skill, and spawnable agentgr-advisorskill — auto-triggers when the user expresses research, video, or content analysis intent without specifying a/grcommand. Prevents suboptimal tool choices (e.g.,/gr:research-deepfor quick questions)gr-advisoragent — sonnet-powered subagent for programmatic workflow routing within agent teams- Agent/Skill/Command conventions in CLAUDE.md — documents required frontmatter fields for plugin contributors
Changed
- CLAUDE.md — added minimal
/gr Plugin Routingsection: recall-first pattern and cost-awareness for/gr:research-deep
0.3.9 - 2026-03-05
Fixed
- AskUserQuestion YAML examples — aligned all 4 code block examples across commands/skills with actual tool schema (
questionsarray withmultiSelectboolean)
0.3.8 - 2026-03-05
Improved
/gr:ingestcommand — now callsknowledge_schemabefore ingesting to discover exact property names; removed hardcoded property table that could drift from schema- video-research skill — added schema-first convention for knowledge ingest workflows
0.3.7 - 2026-03-05
Fixed
- comment-analyst agent — changed model from
haiku(banned) toopusper global policy
0.3.6 - 2026-03-05
Added
knowledge_schematool — returns property names, types, and descriptions for any collection without requiring a Weaviate connection. Call beforeknowledge_ingestto discover expected fields
Improved
knowledge_ingesterror messages — unknown-property errors now include allowedname:typepairs and a hint to callknowledge_schema, eliminating trial-and-error loops
0.3.5 - 2026-03-05
Changed
- 3× faster server startup — lazy-import
google-genaiandweaviateSDKs; deferred from module load to first tool call (fixes Glama Docker build timeout)
0.3.4 - 2026-03-05
Added
__main__.pyentry point — enablespython -m video_research_mcpfor Docker and direct invocation (fixes Glama Docker build)
Fixed
- Stale
__version__—__init__.pynow tracks the actual release version instead of hardcoded0.1.0
0.3.3 - 2026-03-05
Added
- Gemini Deep Research Agent tools — added
research_web,research_web_status,research_web_followup, andresearch_web_cancelfor long-running web-grounded research via the Interactions API - DeepResearchReports knowledge collection — stores completed deep-research reports, usage metadata, and follow-up Q&A; includes cross-references to
ResearchFindingsandWebSearchResults DEEP_RESEARCH_AGENTconfig variable — explicit environment variable and runtime validation for selecting the Interactions API agent/gr:research-deepcommand — interview-driven command workflow for launching and iterating on Deep Research runsresearch-brief-builderskill — brief-quality checklist and challenge templates for high-signal research prompts
Fixed
- Deep Research follow-up tool annotation now correctly marks write behavior (
readOnlyHint=false) - Deep Research launch tracking now evicts stale IDs (TTL + cap) and cleans up terminal interactions to avoid in-memory growth
- Follow-up results are now persisted to Weaviate (
follow_ups_json) instead of storing only follow-up IDs
Changed
researcheragent now includes Deep Research and knowledge-search tools in its default workflow- Installer copy map now ships the new
/gr:research-deepcommand andresearch-brief-builderskill
0.3.2 - 2026-03-03
Added
/gr:getting-startedcommand — interactive first-time setup: verifies config, runs smoke test, shows all available commands and optional features- Installer "Next steps" now links to Gemini API key page and directs users to
/gr:getting-started
Fixed
- Installer: removed unpublished MCP servers —
video-explainer-mcpandvideo-agent-mcpare not on PyPI; the installer was creating broken server entries that failed on startup for every new user - Installer: removed unresolvable env placeholders —
${MLFLOW_TRACKING_URI}was written as a literal string in.mcp.json(no shell expansion); server reads config from~/.config/video-research-mcp/.envinstead
0.3.1 - 2026-03-03
Added
- Local filesystem boundary enforcement — new
local_path_policy.pyvalidates all local file paths againstLOCAL_FILE_ACCESS_ROOT; applied invideo_file,video_batch,research_document_file(PR #41) - Infra mutation auth gating —
infra_configurenow requiresauth_tokenmatchingINFRA_ADMIN_TOKENenv var;infra_cacheread-only ops remain unauthenticated (PR #41) - Prompt injection guardrails — system prompts for content, research, research_document, and knowledge tools now include injection defense instructions (PR #41)
PERMISSION_DENIEDerror category — new error classification inerrors.pyforPermissionError,TimeoutError,httpx.TimeoutException,httpx.NetworkError(PR #41)DocumentPreparationIssuemodel — surfaces file preparation problems inresearch_documentreports (PR #41)- Security tests — adversarial prompt corpus coverage (#43), policy-inheritance guard tests (#44), smoke suite extension (#45)
- 5 new config fields:
research_document_max_sources,research_document_phase_concurrency,local_file_access_root,infra_mutations_enabled,infra_admin_token
Fixed
- Atomic cache writes —
cache.pyandcontext_cache.pynow write via UUID temp files to prevent corruption on concurrent access (PR #41) - Sensitive config redaction —
infra_configureoutput redactsGEMINI_API_KEY,WEAVIATE_API_KEY,COHERE_API_KEY,INFRA_ADMIN_TOKEN(PR #41) - Weaviate setup skill: minor fix in SKILL.md
Changed
url_policy.py— expanded URL validation with stricter enforcementresearch_document.py— concurrency limits via config fields- README: clarify plugin = MCP servers + commands + skills + agents
- Docs: fix tool counts (41 total), review-cycle finalization
0.3.0 - 2026-03-01
Added
- video-agent-mcp — new package: parallel scene generation via Claude Agent SDK (PR #16)
- video-explainer-mcp — new package: 15 tools for synthesizing explainer videos from research (PR #10)
- MLflow tracing —
@tracedecorator on all 24 tools, MLflow MCP server plugin,/gr:tracescommand, and health check in/gr:doctor(PR #12) - Cohere reranking — optional server-side reranking via Cohere when
COHERE_API_KEYis set. Auto-detected; disable withRERANKER_ENABLED=false(PR #18) - Flash summarization — Gemini Flash post-processes search hits to score relevance, generate one-line summaries, and trim unnecessary properties. Disable with
FLASH_SUMMARIZE=false(PR #18) - Media asset pipeline — local file paths propagated through video/content pipelines, shared
gr/mediaasset directory with recall actions (PR #17, #18) generate_json_validated()— dual-path JSON validation inGeminiClient(PR #19, pending)- Project-level git rules — branch protection policy in
.claude/rules/git.md
Changed
- CI: bumped
actions/checkoutv4 to v6,astral-sh/setup-uvv5 to v7,actions/setup-pythonv5 to v6 (PRs #13, #14, #15) - Knowledge search:
rerank_scoreandsummaryfields onKnowledgeHit - Weaviate schema: local media path fields added to collections (PR #17)
Fixed
- MCP transport JSON string deserialization for list parameters in
knowledge_ingestandknowledge_search
Deprecated
knowledge_query— useknowledge_searchinstead, which now includes Cohere reranking and Flash summarization for better results with lower token usage.knowledge_ask(AI-powered Q&A) is unaffected
0.2.0 - 2026-02-28
Added
- Knowledge store — 7 Weaviate collections with write-through storage from every tool
- Knowledge tools —
knowledge_search,knowledge_related,knowledge_stats,knowledge_fetch,knowledge_ingestfor querying stored results - QueryAgent tools —
knowledge_askandknowledge_querypowered byweaviate-agents(optional dependency) - YouTube tools —
video_metadata,video_comments,video_playlistvia YouTube Data API v3 - Context caching — automatic Gemini cache pre-warming after
video_analyzefor both YouTube and local files; session reuse viaensure_session_cache(). Large local files (>=20MB) are context-cached automatically on session creation - Session persistence — optional SQLite backend for video Q&A sessions (
GEMINI_SESSION_DB) - Plugin installer — npm package that copies commands, skills, and agents to
~/.claude/and configures MCP server - MLflow MCP plugin —
/gr:tracescommand for querying, tagging, and evaluating traces;mlflow-tracesskill with field path reference andextract_fieldsdiscipline;mlflow-mcpserver auto-installed viauvx; MLflow health check in/gr:doctor - Diagnostics —
/gr:doctorcommand for MCP wiring, API key, Weaviate, and MLflow connectivity checks - Retry logic — exponential backoff with jitter for Gemini API calls
- Batch analysis —
video_batch_analyzefor concurrent directory-level video processing - PyPI metadata — added classifiers, project URLs, and version alignment with npm
Changed
- Bumped version from 0.1.0 to 0.2.0 (aligned with npm package)
- Unified tool count to 23 across all documentation
0.1.0 - 2026-02-01
Added
- Core server — FastMCP root with 7 mounted sub-servers (stdio transport)
- Video analysis —
video_analyze,video_create_session,video_continue_sessionfor YouTube URLs and local files - Research tools —
research_deep(multi-phase with evidence tiers),research_plan,research_assess_evidence - Content tools —
content_analyze,content_extractwith caller-provided JSON schemas - Search —
web_searchvia Gemini grounding with source citations - Infrastructure —
infra_cache(view/list/clear),infra_configure(runtime model/thinking/temperature) - Structured output — added
GeminiClient.generate_structured()with Pydantic model validation - Thinking support — configurable thinking levels (minimal/low/medium/high) via
ThinkingConfig - Error handling —
make_tool_error()with category, hint, and retryable flag (tools never raise) - Caching — file-based analysis cache with configurable TTL