Skip to content

mustafaham123/ai-ui-cleaner

v0.1.0MIT

Create, review, and repair authored interfaces using retrieved code and visual references.

AI UI Cleaner

AI UI Cleaner is a reference-grounded interface creation, review, and repair plugin for GPT/Codex-compatible hosts and Claude Code. It combines:

  • A portable Agent Skill that directs Figma and code-based creation, redesign, implementation, and system-level review.
  • An MCP server exposing focused reference-search tools.
  • A local-first hybrid RAG layer with keyword relevance, hashed-vector similarity, metadata boosts, filtering, and diversity reranking.
  • Conservative provenance and code-reuse gates.

The committed seed records are original abstract design patterns. Harvested reference data, code, source snapshots, and preview images belong outside Git. The Cloudflare backend and bounded collector are implemented; see the cloud corpus setup for deployment, existing collection coverage, and imports.

Quick start

npm install
npm run build
npm run check

Start the local stdio server:

npm start

Start a stateless streamable HTTP endpoint:

npm run start:http

The endpoint is available at http://localhost:3000/mcp; health information is at http://localhost:3000/health.

Plugin layouts

  • plugin.json, mcp.json, and skills/ form the portable Agent Plugins package.
  • .codex-plugin/plugin.json is the Codex compatibility manifest.
  • .claude-plugin/plugin.json and .mcp.json support Claude-compatible installation.
  • Both local manifests start the committed, bundled dist/index.cjs, so plugin installs do not need production dependencies. Run npm run build whenever server source changes.

For ChatGPT web usage, deploy the server in HTTP mode to a stable HTTPS endpoint and change the MCP entry to:

{
  "$schema": "https://agent-plugins.org/schemas/1.0.0/mcp.schema.json",
  "mcpServers": {
    "ai_ui_cleaner": {
      "type": "streamable-http",
      "url": "https://your-domain.example/mcp"
    }
  }
}

Do not commit credentials to either MCP configuration.

MCP tools

ToolPurpose
search_referencesHybrid search with design and license filters; never returns raw code.
get_referenceRetrieves one selected reference and its provenance.
get_reference_assetFetches one allowlisted, curated screenshot with MIME, size, redirect, and digest checks.
get_code_assetReturns code only when reuse is licensed and curator review is complete.
reference_statsShows corpus coverage and reusable-code counts.

All model-facing tools are read-only. New records are added through local ingestion or the separate authenticated cloud admin API.

Add reference data

Prepare JSON or JSONL using examples/reference.example.json, then run:

npm run ingest -- path/to/export.jsonl

This writes to data/local/references.jsonl, which is intentionally gitignored. The server reloads changed data automatically. Set AI_UI_CLEANER_DATA_PATHS to a platform-delimited list of other corpus files when required.

Imported code is always downgraded to unreviewed by default. After manually verifying source, license, dependencies, security, and attribution, a curator can set reviewStatus to reviewed and re-import with:

npm run ingest -- reviewed.jsonl --preserve-review-status

The following fields are important for retrieval quality:

  • pageTypes, industries, moods, and components
  • summary, whyItWorks, and avoidWhen
  • implementation: general approach, pattern-specific steps, responsive/accessibility notes, and adaptation instructions
  • source.name, source.url, author, capture date, and license
  • Technology and framework requirements for reusable code

Normalize exports from 21st.dev, CodePen, Figma, Dribbble, recent.design, SaaS galleries, or another permitted source into this schema. Store screenshots on controlled object storage and add their URL, dimensions, media type, alt text, and SHA-256 to assets. Use official APIs or authorized exports where available. Respect robots rules, terms, access controls, attribution, and asset/code licenses; a publicly viewable page does not automatically permit republication or code reuse.

Production corpus architecture

Do not commit the full database, screenshots, generated embeddings, or scraped code archives to GitHub. Keep the repository limited to application code, the skill, schemas, migrations, small original seed records, and test fixtures.

Use this production split:

LayerStoreContents
RepositoryGitHubMCP code, skill, schemas, migrations, tiny seeds, tests
Metadata and retrievalCloudflare D1 + Vectorize (implemented)Reference metadata, normalized text, tags, license state, embeddings, searchable reviewed-code chunks
Binary assetsCloudflare R2 (implemented)Preview images, Figma exports, permitted snapshots, license notices
Ingestion workersJob runner or queueAuthorized fetching, screenshot capture, normalization, deduplication, embeddings, license review
MCP runtimeContainer or serverless serviceSearch, record retrieval, signed asset access, and code-license enforcement

The normal flow is:

authorized source/export
        ↓
ingestion worker → object storage for images and archives
        ↓
normalized metadata + embeddings → database/search index
        ↓
MCP tools → selected records, screenshots, and reviewed code
        ↓
AI UI Cleaner skill → design or review workflow

Object-storage URLs should use a controlled hostname listed in AI_UI_CLEANER_ASSET_HOSTS. Prefer short-lived signed URLs for private assets. Do not use Git LFS as the primary corpus: it helps with a few large versioned fixtures, but a growing retrieval database still makes clones, history, updates, and queries inefficient.

Retrieval behavior

The built-in retriever requires no embedding API key. It combines BM25-style term scoring with a deterministic hashed token/trigram vector, then applies explicit metadata boosts and source/type diversity penalties. This is suitable for local development and predictable testing.

The Cloudflare backend uses D1 FTS5 plus Workers AI embeddings and Vectorize, with rank fusion and source/type diversity. Set AI_UI_CLEANER_CLOUD_URL and MCP_READ_TOKEN to make the stdio server forward to the hosted corpus. Local emulation runs FTS5 without the semantic service. See deployment and testing.

Environment

VariableDefaultMeaning
MCP_TRANSPORTstdioSet to http for streamable HTTP.
MCP_HOST127.0.0.1HTTP bind address; containers usually use 0.0.0.0.
MCP_ALLOWED_HOSTSlocalhost validationOptional comma-separated HTTP Host allowlist.
PORT3000HTTP listening port.
AI_UI_CLEANER_DATA_PATHSseed + local JSONLPlatform-delimited local corpus paths.
AI_UI_CLEANER_ASSET_HOSTSnoneComma-separated hostnames permitted for screenshot retrieval.
AI_UI_CLEANER_MAX_ASSET_BYTES8388608Maximum fetched screenshot size.

Security model

  • External narrative fields are plain text, HTML-stripped, and checked for common prompt-injection phrases during ingestion.
  • Source URLs are restricted to HTTP and HTTPS.
  • Raw code never appears in ordinary search or reference results.
  • get_code_asset requires both license.reuseAllowed: true and reviewStatus: reviewed.
  • Scraped code is never executed by ingestion or retrieval.
  • MCP results explicitly tell the model to treat retrieved content as untrusted evidence.