k-dense-ai/reinforcement-learning-researcher
v1.1.0MIT
Reasons from MDP/POMDP structure, Bellman contraction and the deadly triad, and policy-gradient variance through Gymnasium 1.x/MuJoCo v5/ALE v5 protocols, CleanRL/SB3/JAX (MJX, MuJoCo Playground) stacks, rliable IQM with stratified bootstrap CIs, Minari offline datasets, and GRPO/RLVR post-training while treating truncation-as-termination bootstrap bugs, seed and hyperparameter selection bias, reward hacking, and offline extrapolation error as first-class failure modes.
What this package declares
The file a client reads when it loads this plugin, exactly as this revision carries it.
plugin.json
{
"$schema": "https://agent-plugins.org/schemas/1.0.0/plugin.schema.json",
"name": "reinforcement-learning-researcher",
"version": "1.1.0",
"description": "Reasons from MDP/POMDP structure, Bellman contraction and the deadly triad, and policy-gradient variance through Gymnasium 1.x/MuJoCo v5/ALE v5 protocols, CleanRL/SB3/JAX (MJX, MuJoCo Playground) stacks, rliable IQM with stratified bootstrap CIs, Minari offline datasets, and GRPO/RLVR post-training while treating truncation-as-termination bootstrap bugs, seed and hyperparameter selection bias, reward hacking, and offline extrapolation error as first-class failure modes.",
"author": {
"name": "K-Dense",
"url": "https://www.k-dense.ai"
},
"homepage": "https://github.com/K-Dense-AI/scientific-agents",
"repository": "https://github.com/K-Dense-AI/scientific-agents",
"license": "MIT",
"keywords": [
"science",
"agents-md",
"expert-profile",
"reinforcement-learning-researcher"
]
}