자동화된 ML 연구를 위한 ARIS 프로젝트
🚀 프로젝트 소개
ARIS(Automatic Research In Sleep)는 경량의 Markdown 기반 도구로, 자율적인 머신러닝(ML) 연구를 지원합니다. 이 프로젝트는 다양한 모델 간의 협업을 통해 아이디어 발견과 실험 자동화를 가능하게 하여, 연구자가 잠자는 동안에도 연구를 진행할 수 있도록 돕습니다.
✨ 주요 기능
- Markdown 파일을 사용한 경량 구조로, 의존성 없음
- 다양한 LLM(대형 언어 모델)과의 호환성 제공
- 자동화된 연구 프로세스와 실험 실행 기능
🛠️ 기술 스택
이 프로젝트는 Python으로 개발되었으며, Claude Code, Codex, OpenClaw 등 다양한 LLM 에이전트와의 통합을 지원합니다. 또한, 별도의 프레임워크나 데이터베이스 없이 작동합니다.
💡 활용 방법
개발자는 ARIS를 통해 자신의 연구 워크플로우를 자동화하고, 다양한 LLM을 활용하여 연구의 효율성을 높일 수 있습니다. 이 시스템은 사용자가 원하는 대로 수정하고 확장할 수 있는 유연성을 제공합니다.
📄 Original (English)
ARIS ⚔️ (Auto-Research-In-Sleep) — Lightweight Markdown-only skills for autonomous ML research: cross-model review loops, idea discovery, and experiment automation. No framework, no lock-in — works with Claude Code, Codex, OpenClaw, or any LLM agent.
Auto-claude-code-research-in-sleep (ARIS ⚔️🌙)
中文版 README | English
🌙 Let Claude Code do research while you sleep. Wake up to find your paper scored, weaknesses identified, experiments run, and narrative rewritten — autonomously.
🪶 Radically lightweight — zero dependencies, zero lock-in. The entire system is plain Markdown files. No framework to learn, no database to maintain, no Docker to configure, no daemon to babysit. Every skill is a single
SKILL.mdreadable by any LLM — swap Claude Code for Codex CLI, OpenClaw, Cursor, Trae, Antigravity, Windsurf, or your own agent and the workflows still work. Fork it, rewrite it, adapt it to your stack.💡 ARIS is a methodology, not a platform. What matters is the research workflow — take it wherever you go. 🌱
Custom Claude Code skills for autonomous ML research workflows. These skills orchestrate cross-model collaboration — Claude Code drives the research while an external LLM (via Codex MCP) acts as a critical reviewer. 🔀 Also supports alternative model combinations (Kimi, LongCat, DeepSeek, etc.) — no Claude or OpenAI API required. For example, MiniMax-M2.7 + GLM-5 or GLM-5 + MiniMax-M2.7. 🤖 Codex CLI native — full skill set also available for OpenAI Codex. 🖱️ Cursor — works in Cursor too. 🖥️ Trae — ByteDance AI IDE. 🚀 Antigravity — Google's agent-first IDE. 🆓 Free tier via ModelScope — zero cost, zero lock-in.
💭 Why not self-play with a single model? Using Claude Code subagents or agent teams for both execution and review is technically possible, but tends to fall into local minima — the same model reviewing its own patterns creates blind spots.
Think of it like adversarial vs. stochastic bandits: a single model self-reviewing is the stochastic case (predictable reward noise), while cross-model review is adversarial (the reviewer actively probes weaknesses the executor didn't anticipate) — and adversarial bandits are fundamentally harder to game.
💭 Why two models, not more? Two is the minimum needed to break self-play blind spots, and 2-player games converge to Nash equilibrium far more efficiently than n-player ones. Adding more reviewers increases API cost and coordination overhead with diminishing returns — the biggest gain is going from 1→2, not 2→4.
Claude Code's strength is fast, fluid execution; Codex (GPT-5.4 xhigh) is slower but more deliberate and rigorous in critique. These complementary styles — speed × rigor — produce better outcomes than either model talking to itself.
MIT License