빛의 속도로 작동하는 LLM 추론 엔진
🚀 프로젝트 소개
TokenSpeed는 에이전트 작업을 위한 빛의 속도로 작동하는 LLM(Large Language Model) 추론 엔진입니다. TensorRT-LLM 수준의 성능과 vLLM 수준의 사용성을 제공하여, 생산 환경에서 가장 효율적인 추론 엔진을 목표로 하고 있습니다.
✨ 주요 기능
- 모델링 레이어: 사용자 정의 병렬 처리 로직 없이도 효율적인 통신을 생성
- 스케줄러: C++ 제어 평면과 Python 실행 평면으로 구성된 요청 생애 주기 관리
- 플러그 가능한 커널 시스템: 빠른 MLA(Multi-head Latent Attention) 구현 포함
🛠️ 기술 스택
이 프로젝트는 Python으로 개발되었으며, C++와 TensorRT를 활용하여 성능을 극대화합니다. 또한, 다양한 런타임 기능과 플랫폼 최적화를 지원합니다.
💡 활용 방법
개발자는 TokenSpeed를 사용하여 고성능 LLM 추론을 구현하고, 다양한 모델을 지원하는 서버를 쉽게 구축할 수 있습니다. 문서화된 가이드를 통해 손쉽게 시작할 수 있습니다.
📄 Original (English)
About
TokenSpeed is a speed-of-light LLM inference engine.
README
TokenSpeed is a speed-of-light LLM inference engine designed for agentic workloads, with TensorRT-LLM-level performance and vLLM-level usability. Our goal is to be the most performant inference engine for production agentic workloads.
Core components:
- Modeling layer: local-SPMD design with a static compiler that generates collective communication from module-boundary placement annotations, so users do not hand-write parallelism logic.
- Scheduler: C++ control plane and Python execution plane. Request lifecycle, KV cache ownership, and overlap timing are encoded as a finite-state machine, with safe KV resource reuse enforced by the type system at compile time.
- Kernels: pluggable, layered kernel system with a portable public API and a centralized registry including one of the fastest MLA (Multi-head Latent Attention) implementations on Blackwell for agentic workload.
- Entrypoint: SMG-integrated AsyncLLM for low-overhead CPU-side request handling.
Performance Comparison
License
MIT License
|
댓글을 작성하시려면 로그인이 필요합니다.