최신 LLM 구축을 위한 가이드
🚀 프로젝트 소개
이 프로젝트는 최신 언어 모델(LLM)을 처음부터 끝까지 구축하는 방법을 안내하는 12장 구성의 인터랙티브 교과서입니다. 각 코드 라인에 주석이 달려 있어 초보자도 쉽게 이해할 수 있도록 설명합니다.
✨ 주요 기능
- 3,900줄 이상의 주석이 달린 코드로 구성
- Transformer 모델의 내부 작동 원리를 심도 있게 설명
- 기초 Python 지식만으로도 접근 가능
🛠️ 기술 스택
주요 기술로는 Python, PyTorch, Jupyter Notebook이 사용되며, Attention Mechanism과 Transformer 아키텍처에 대한 이해를 돕습니다.
💡 활용 방법
개발자는 이 프로젝트를 통해 LLM의 기초부터 심화까지 체계적으로 학습하고, 실제 모델을 구축하는 경험을 쌓을 수 있습니다.
📄 Original (English)
Build a modern LLM from scratch. Every line commented. Explained like we are five.
🧠 How to Train Your GPT
A guide to building a world-class language model from absolute scratch. Taught like you're five. Built like you're an engineer.
I made this with the goal of learning something I didn't understand completely. Specifically the attention part. I use AI a lot to understand key concepts and verifying them.
📖 What Is This?
This is a 12-chapter, 3,900+ line interactive textbook that teaches you how to build, train and run a modern language model from absolute scratch. The same family of architecture behind ChatGPT, Claude, LLaMA and Mistral.
You won't just read about Transformers. You'll write every line yourself: tokenizer, embeddings, attention, training loop, inference engine. Every single line annotated to explain what it does and why it's there.
🤔 Why This Exists
Most ML tutorials fall into one of two traps:
| ❌ Too Shallow | ❌ Too Academic | ✅ This Guide |
|---|---|---|
model = GPT().fit(data) |
40-page papers, dense notation | 5-year-old analogies → full working code |
MIT License