Skip to content

furox-art/eq-layer

v0.2.0MIT

Inspectable intent, affect, repair, uncertainty, and factored-action control skill for general-purpose AI agents.

eq-layer

A control layer that gives an LLM a decision about how to speak before it speaks — instead of asking it to be more empathetic in a prompt.

Why this exists

Asking a model to be more empathetic is style transfer. It produces more of the same flat warmth, which reads as uncanny rather than attentive. The failure modes are structural, not knowledge gaps:

  • Flat affect. "That sounds really tough!" arrives at the same register regardless of what was said.
  • Escalation blindness. The user is frustrated, then angry, then threatening. The model holds one tone for all three turns.
  • Validation without motion. The feeling is acknowledged; the next step never comes.
  • Repair collapse. After a correction the model either capitulates entirely or starts defending itself. There is no middle.

So this repo does not retrain the model. It adds a layer:

  1. Affect state — {valence, arousal, escalation_delta, stance, subtext} as an explicit intermediate representation, tracked across turns. escalation_delta is the point: snapshot classification cannot see a trend.
  2. Intent belief state — the learned classifier preserves the full posterior over action_request/status_check/explanation/question/statement/unknown instead of collapsing uncertainty to one confidence number. A transparent one-step Bayes-risk router chooses whether to act or clarify from that distribution. This is POMDP-inspired uncertainty handling, not a full POMDP.
  3. Conversation health state — repair identifies observable user→assistant misalignment and its target; interaction_quality tracks structural conversation failure separately from emotion using current score, delta, repeated failures, unresolved repair, and clarification load.
  4. Policy selection — affect and intent jointly choose a response-policy family (mirror, direct, repair, hold, boundary, ...). Selection is a discrete decision, not generation.
  5. Factored action compilation — the selected policy is split into orthogonal controls: task_move, social_move, repair_move, and realization. The task is driven primarily by user intent, so affective adaptation cannot silently replace what the user actually asked for.
  6. Steering — the factored action, source policy, repair state, interaction quality, and canonical request are injected into the decode path so the control decision actually lands in the tokens.
  7. Structural response audit — the generated response is checked against inspectable control contracts such as question budget, required clarification, unsupported privacy promises, unknown-stance admissions, repair repetition, and low-verbosity overruns. This is not a factual truth checker.

Prior art and novelty boundary

EQ-Layer does not claim that affect-aware dialogue policy, user modelling, repair, intent routing, or adaptive dialogue management are new. Those ideas have substantial pre-LLM precedent in TRAINS/grounding work, adaptive TOOT, RavenClaw, user-tailored generation, affect-sensitive AutoTutor, affective POMDP dialogue management, and related systems.

The defensible research question is narrower: whether an explicit, model-agnostic control layer can improve a fixed general-purpose language model by separating intent, affect, conversational evidence, policy selection, and generation steering while keeping those intermediate decisions inspectable.

See docs/PRIOR_ART_AND_NOVELTY.md for the historical map, claims to avoid, and architecture lessons carried forward into this project.

Status

Early. The harness runs; the numbers are development signals, not a real-world EQ benchmark.

There is no standard benchmark for EQ. eval/ is a first attempt at one, not a finished measurement. If you improve the scorer, that contribution matters as much as adding a policy.

Agent Skills: Codex, Claude, Copilot, Cursor, Antigravity, and Grok

EQ-Layer is packaged as a portable Agent Skill in skills/eq-layer/ and as an Agent Plugins 1.0 package via plugin.json. Repository discovery shims are included for the major agent ecosystems:

AgentRepository path
OpenAI Codex.codex/skills/eq-layer/
Anthropic Claude Code.claude/skills/eq-layer/
GitHub Copilot.github/skills/eq-layer/
Cursor.cursor/skills/eq-layer/
Google Antigravity.agents/skills/eq-layer/
xAI Grok Build.grok/skills/eq-layer/

For a user-level install across all supported discovery roots:

python tools/install_agent_skills.py --target all

The skill includes a local routing script and reports whether it used the lightweight portable path or the full learned research pipeline.

printf '%s' '[{"role":"user","content":"devam et"}]' \
  | python skills/eq-layer/scripts/route.py --pretty

See skills/eq-layer/references/VENDOR_COMPATIBILITY.md for verified discovery paths and source documentation.

Quick start

Install the learned intent tracker:

pip install "eq-layer[ml]"
from eq_layer import ConversationTracker, Selector, Steer, TrainedAffect, TrainedIntent, TrainedSubtext, compose_action

messages = [
    {"role": "user", "content": "Randevumu üç kez değiştirdiler"},
]
affect = TrainedAffect.load("artifacts/emobank_affect.joblib")
subtext = TrainedSubtext.load("artifacts/xdailydialog_subtext.joblib")
tracker = ConversationTracker(affect=affect, subtext=subtext)
tracked = tracker.infer(messages)
state = tracked.state
intent = TrainedIntent.from_bundled().infer(messages)
selection = Selector().select(state, intent)
action = compose_action(
    selection.policy,
    intent,
    repair=tracked.repair,
    interaction_quality=tracked.interaction_quality,
)
print(selection.policy.name, action.to_dict())

prompt = Steer.build(
    selection.policy,
    intent,
    repair=tracked.repair,
    interaction_quality=tracked.interaction_quality,
).apply_to_prompt(messages[-1]["content"])

Run the seeded cases. Exits non-zero on any mismatch or any policy that no case can reach:

python eval/run.py

Run the intent benchmark separately:

python eval/intent_eval.py

It compares the learned adapter with the heuristic fallback and reports five-fold cross-validation on the bundled corpus.

Train the affect regressor from the pinned official EmoBank release:

python tools/train_emobank_affect.py --output artifacts/emobank_affect.joblib
python eval/emobank_affect_eval.py

EmoBank is not bundled into this repository. The training/evaluation path uses the upstream train/dev/test split and records its pinned commit and CC-BY-SA-4.0 provenance in THIRD_PARTY_DATA.md.

Train the learned dialogue-signal tracker from pinned XDailyDialog files:

python tools/train_xdailydialog_subtext.py --output artifacts/xdailydialog_subtext.joblib
python eval/xdailydialog_subtext_eval.py

XDailyDialog provides supervised dialogue-act and basic-emotion labels. EQ-Layer does not pretend that challenge, demand, or exhaustion are direct upstream labels: those are conservative compositions of learned dialogue signals with learned V/A state. See THIRD_PARTY_DATA.md for the license-chain caution on the English data.

Intent is a control signal, not mind-reading

The default intent path is now learned rather than keyword-selected. TrainedIntent fits a balanced logistic-regression classifier over combined word and character TF-IDF features. The bundled corpus currently contains Turkish and English examples for action_request, status_check, explanation, question, statement, and unknown.

The classifier emits a full posterior distribution, exposed as IntentBelief. Production routing no longer discards that distribution behind a single confidence/margin threshold. decide_intent_action computes the expected interaction loss of each response move and clarification, then chooses the minimum-risk action. For example, ambiguity between a status check and a generic question can still be answered because the semantic mismatch cost is small, while ambiguity between executing an action and merely responding to a statement can trigger clarification.

The current loss matrix is hand-specified and transparent, not learned or claimed optimal. It encodes relative interaction costs. EQ-Layer therefore audits routing sensitivity rather than treating one matrix as ground truth: eval/intent_risk_sensitivity.py sweeps clarification cost and performs one-at-a-time ±20% perturbations of every non-zero loss cell, reporting action flips overall and by source stratum. This is a local robustness diagnostic, not loss-matrix optimization, and it must not be tuned against blinded human preference outcomes. The old confidence and margin fields remain serialized for backwards compatibility and diagnostics; they no longer choose the learned intent route.

Intent-risk sensitivity audit

The production control path has also been run over all 120 external held-out cases without generating LLM responses. Under one-at-a-time ±20% perturbations of every non-zero loss cell, 7,332 / 7,440 routing decisions were unchanged (98.55% local decision stability).

That result does not validate the loss matrix. Scaling the entire clarification-cost row changes routing substantially:

  • production clarification rate: 45.0%
  • clarification cost ×0.75: 69.17%
  • clarification cost ×1.25: 19.17%
  • clarification cost ×1.50: 9.17%

The task_repair stratum is especially uncertain: 26 / 30 cases route to clarification (86.67%) with mean fused intent entropy 0.8954. This may reflect genuine ambiguity in repair turns, but it is not used to tune the matrix. The current loss table remains hand-specified and uncalibrated.

Machine-readable audit: eval/results/intent_risk_heldout_120_audit.json.

Explicit constraints such as brief, scope_limited, and no_guess remain deterministic because they are control requirements, not semantic labels.

HeuristicIntent remains available as a zero-dependency fallback. It does not manufacture a learned posterior and is no longer the adapter used by the main evaluation harness.

Affect can still override intent during sustained escalation. A clear action request should drive the response in a calm turn; it should not erase a multi-turn escalation signal.

The bundled data is small and synthetic. eval/intent_eval.py therefore reports both a fixed held-out score and five-fold cross-validation; neither is presented as evidence of production-level semantic understanding.

Verified development baselines

The current CI-verified EmoBank regression baseline uses the official upstream train/dev/test split (8062 / 1000 / 1000). On the untouched test split:

dimensionMAEtrain-mean baseline MAESpearman rho
Valence0.21250.24490.5460
Arousal0.17820.19040.3120
Dominance0.14940.15630.2573

The regressor beats the train-mean baseline on MAE for all three dimensions, but Arousal and Dominance rank correlations are still modest. The machine-readable record is eval/results/emobank_affect_baseline.json.

Dialogue-signal baseline

The learned dialogue-signal tracker was verified on the pinned XDailyDialog English train/dev/test files (83035 / 8025 / 7716 utterances). On the untouched test file:

taskaccuracymacro-F1majority accuracy
dialogue act0.80250.72740.4557
basic emotion0.82490.43100.8166

The dialogue-act signal is substantially above the majority baseline. The emotion signal is much weaker: macro-F1 is only 0.431 and accuracy is only slightly above the majority baseline on test (and below it on dev). EQ-Layer therefore treats emotion as supporting evidence rather than standalone proof of subtext. The machine-readable result is eval/results/xdailydialog_signal_baseline.json.

Higher-level subtext audits

Direct higher-level supervision was investigated rather than assumed:

source / targetverified resultrouting decision
DialogBank correction4 direct examplesNO-GO — insufficient supervision
DialogBank disagreement1 direct exampleNO-GO — insufficient supervision
Coarse Discourse disagreementtest ROC-AUC 0.8111, AP 0.1420, F1 0.2278NO-GO — ranking signal exists, but precision/recall are insufficient for challenge routing
DBDC3 hard breakdowntest ROC-AUC 0.5057, AP 0.2555NO-GO — effectively chance-level text-only generalisation

These results are preserved rather than optimized away. The corresponding machine-readable records are eval/results/dialogbank_label_inventory.json, eval/results/coarse_disagreement_no_go.json, and eval/results/dbdc3_breakdown_no_go.json.

Layer compute cost and decision quality

Component metrics do not say what the layer costs to run. A paired A/B run of the same 15 seeded requests with the layer enabled and disabled, 100 passes per arm, measured the layer's own deterministic compute with no LLM in either arm:

armadded median (ms)added p95 (ms)intent holdout accuracy (36)appropriateness rubric
learned intent adapter1.16121.46361.0000.9917
heuristic zero-dependency0.11040.22650.27781.0000

Classification and routing are timed independently rather than lumped: with the learned adapter, intent inference is ~92% of classification and routing is ~2% of the total. The two adapters are not interchangeable — both score 1.000 on the 4 labelled seed cases, but only the learned adapter holds up on the 36-case holdout. Latency is machine-dependent and moved ~9% between runs on the same machine, so treat the learned figure as approximately 1.2 ms median rather than an exact constant.

The measured quantity is the layer's compute cost and the structural quality of its decision, not response quality: there is no model in the measured path to generate prose. End-to-end quality still needs the paired human evaluation in eval/AB_PROTOCOL.md.

python eval/ab_latency_eval.py

Full methodology, tables and run-to-run variation are in docs/AB_LATENCY_BENCHMARK.md. Machine-readable record: eval/results/ab-latency.json.

Evidence provenance

A subtext label is not enough by itself. SubtextDecision.evidence_level records how the label was obtained:

  • verified — caller/external annotation supplied the state.
  • direct_learned — the upstream training target directly matches the emitted label.
  • derived — multiple learned signals are composed into a higher-level control label.
  • structural — explicit conversational form or marker triggered the label.
  • fallback — learned evidence was too weak and an auditable fallback was used.

For example, XDailyDialog directly supervises question, but it does not directly supervise EQ-Layer's challenge or exhaustion; those remain derived. StanceDecision uses the same principle: factual user_right/user_wrong stays unknown unless externally verified.

Blind response-level evaluation

Component metrics do not establish that EQ-Layer improves the final response. The repository therefore includes a condition-blind paired human-evaluation harness. The complete pre-registered protocol is in eval/AB_PROTOCOL.md.

Generate paired responses with the same base model in both arms:

python eval/ab_generate.py heldout.jsonl pairs.jsonl manifest.json \
  --model-command "python model_adapter.py" \
  --model-id SAME_BASE_MODEL_VERSION \
  --affect-model artifacts/emobank_affect.joblib \
  --subtext-model artifacts/xdailydialog_subtext.joblib \
  --temperature 0 \
  --seed 42 \
  --git-commit <eq-layer-commit>

The configured command receives JSON on stdin and returns either plain response text or JSON containing text/response. Secrets belong in environment variables, not command arguments. eval/cases.jsonl is rejected by default because it is a development set, and gold annotations are ignored unless an explicit oracle-analysis flag is supplied.

Prepare pairs as JSONL:

{"id":"case-001","context":[{"role":"user","content":"..."}],"baseline":"...","eq":"..."}

Blind and randomize them:

python eval/ab_prepare.py pairs.jsonl ballot.jsonl key.json --seed 42

Keep key.json away from raters. For every case, rate A/B/tie on:

  • intent_fidelity
  • appropriateness
  • actionability
  • non_patronizing
  • overall

Then score a single rater:

python eval/ab_score.py rated_ballot.jsonl key.json

Or score multiple raters on the same blinded cases:

python eval/ab_score_multi.py key.json rater1.jsonl rater2.jsonl rater3.jsonl

The single-rater scorer reports EQ wins/losses/ties, non-tie win rate, Wilson 95% confidence intervals, and an exact two-sided sign test. The multi-rater scorer additionally aggregates case-level majority preference and reports Fleiss' kappa for inter-rater agreement; it does not treat every rater vote as an independent case.

The harness does not manufacture a result: response generation and human ratings must come from a real pre-registered comparison.

Passive response auditor

A generated EQ response is also inspected by ResponseAuditor. The A/B generation pipeline records the result under generation.response_audit and the manifest records response_auditor_mode=passive-metadata-only.

The auditor currently checks only structurally observable control violations:

  • empty response;
  • question-budget overrun;
  • missing required clarification;
  • unsupported absolute privacy promise;
  • stance=unknown combined with an explicit "you are right / I was wrong" admission;
  • near-duplicate recent assistant responses;
  • repeated prior answer while a repair is active;
  • large low-verbosity overruns.

It does not decide whether factual claims are true, whether the user's position is correct, or whether advice is substantively good. Hard audit failures currently request regeneration in metadata only; automatic regeneration is deliberately disabled in the A/B experiment so the evaluation protocol is not silently changed.

Selection

Policies declare preconditions, and the most constrained applicable policy wins. There is no first-match rule, so adding a policy cannot silently change what a looser one does to a state you did not think about.

Two consequences worth arguing about:

  • Implicit tie-breaking is forbidden. Selection ranks by specificity, then explicit priority. If two applicable policies still tie, the selector raises an ambiguity error instead of silently picking whichever was declared first.
  • stance is not guessed. Whether the user is right is a semantic judgement, so the tracker returns unknown unless a case or caller supplies one. Policies that need it declare user_is_right or stance_unknown, so "I could not tell" stays visible in the output instead of becoming a confident wrong answer.

Cases can carry annotated fields (stance, subtext) for signals that keywords genuinely cannot recover — exhaustion is the current example.

Layout

eq_layer/
  policies.py      joint affect/intent policy taxonomy + selector
  actions.py       task/social/repair/realization factorization
  auditor.py       structural post-generation control-contract audit
  interaction.py   structural repair lifecycle + interaction-quality state
  affect.py           affect-state tracker interface + zero-dependency structural fallback
  trained_affect.py   learned VAD regressor + multi-turn affect adapter
  trained_subtext.py     learned dialogue-act/emotion signals + conservative subtext derivation
  trained_disagreement.py experimental direct-disagreement model; verified no-go for routing
  trained_breakdown.py   experimental text-only breakdown model; verified no-go for routing
  tracker.py             composes affect, subtext, stance, repair and interaction quality
  intent.py           intent state + zero-dependency fallback
  intent_belief.py    posterior belief + transparent Bayes-risk router
  trained_intent.py   learned TF-IDF + logistic-regression posterior adapter
  data/intent_train.jsonl  bundled training corpus
  steer.py               policy + intent → decode path, and the scorer
  response_experiment.py same-model baseline-vs-EQ generation engine
eval/
  cases.jsonl          development cases, not final A/B evidence
  AB_PROTOCOL.md       pre-registered response-level evaluation protocol
  ab_generate.py       paired response generation
  ab_prepare.py        A/B blinding
  ab_score.py          single-rater scoring
  ab_score_multi.py    multi-rater scoring
  ab_latency_eval.py   layer-on vs layer-off latency + decision-quality A/B
  run.py

Scoring

Three metrics, deliberately narrow:

  • specificity — does the response name the concrete thing the user actually said, or a category of difficulty?
  • genericness — count of empty validation phrases. Penalised on purpose: these are cheap under human raters, so models learn to produce them unless a cost is attached.
  • escalation latency — how many turns pass after tension rises before the register changes.

If genericness is not penalised, a model will score well by saying nothing in particular.

Affect and conversational signals are now partly learned

TrainedAffect learns continuous Valence, Arousal and Dominance from EmoBank, then converts Valence to [-1, 1] and Arousal to [0, 1] for policy state. Multi-turn escalation_delta is computed from consecutive learned arousal predictions rather than keyword counts.

TrainedSubtext separately learns dialogue-act and basic-emotion signals from XDailyDialog. Those supervised signals are then combined with V/A state to derive higher-level control labels conservatively:

  • learned question -> question
  • learned directive + independent heat -> demand
  • learned question + anger/disgust + high arousal -> challenge
  • learned sadness + negative low-arousal V/A -> exhaustion / resignation

correction and disclosure boundaries remain structural because the chosen upstream data does not directly supervise those categories.

stance=user_right/user_wrong is not inferred from generic dialogue. Whether a user is factually or procedurally correct requires external evidence. AnnotationStanceResolver therefore emits unknown unless a verifier or annotation supplies a stance. This is intentional rather than a missing confidence threshold.

The remaining limitations are:

  • escalating_past_n excludes the current turn. The question is whether the preceding turns were rising, so one polite message after three hostile ones does not reset the register.
  • question vs challenge turns on whether escalation markers are present, not on punctuation. Sözleşme kaç gün geçerli? is a question; Sen de mi?! is a challenge.
  • Exhaustion/resignation still lacks direct supervision. The current label is a derived composition of learned sadness plus low-valence/low-arousal state. It must not be described as a directly learned exhaustion classifier.
  • First-word matching strips punctuation. Neden? must match neden or the one-word repeated question — the exact shape escalating exists to catch — silently never fires.

The remaining hard gap is direct, transferable supervision for higher-level states such as correction, exhaustion/resignation, and factual stance. Two plausible direct proxies — disagreement detection and text-only breakdown detection — were tested and retained as no-go results instead of being wired into policy routing. The tracker exposes where labels are direct, derived, structural, fallback, or externally verified.

Contributing

Cases are the main entry point. Open one for a failure mode you can describe concretely — especially the ones where two policies seem equally right.

Policy changes go through an RFC first: .github/RFC_TEMPLATE.md.

License

MIT. See LICENSE.