GitHub 트렌딩을 그대로 나열하지 않고, Claude Code · RAG · 로컬 AI · 에이전트 워크플로우 · 평가 도구 · AI 앱 빌더 관점에서 실제 빌더가 쓸 만한 오픈소스 스택을 다시 정리합니다.
AI 엔지니어링의 기본부터 실제 서비스 구축까지 다루는 실습 중심의 학습 자료다. 에이전트와 컴퓨터 비전 등 다양한 AI 분야를 포함한다.
Learn it. Build it. Ship it for others.
PDF·DOCX·PPTX·이미지·오디오 등 모든 걸 LLM에 먹이기 좋은 깨끗한 마크다운으로 변환. RAG 전처리에서 압도적으로 편함.
Convert anything (PDF, Word, PPTX, images, audio) into clean Markdown for LLM ingestion.
Python으로 고성능 웹 애플리케이션을 구축하기 위한 경량 비동기 웹 프레임워크이다. 효율적인 아키텍처로 확장성과 성능을 동시에 잡을 수 있다.
Non-autoregressive System 1 decision engine. Typed choice, score and yes/no decisions over any text in a single forward pass, in 100+ languages, with a router that picks the right checkpoint per request.
OSS LLM 추론 엔진의 사실상 표준. PagedAttention 기반으로 처리량이 매우 높아요.
A high-throughput and memory-efficient inference and serving engine for LLMs.
Pandas보다 5~30배 빠른 dataframe. lazy evaluation + Arrow 기반으로 메모리 효율도 좋음.
Dataframes powered by a multithreaded, vectorized query engine, written in Rust.
Pandas + DuckDB 조합이 데이터 분석의 새 표준. SQL로 거대 파일 즉석 쿼리.
DuckDB — an in-process SQL OLAP database management system.
MS의 그래프 기반 RAG. 일반 RAG보다 multi-hop 추론이 강해서 복잡 도메인에 적합.
A modular graph-based Retrieval-Augmented Generation system.
Hugging Face의 모델 허브를 다루는 표준 라이브러리. 새 모델이 나오면 가장 먼저 여기에 들어와요.
State-of-the-art Machine Learning for Pytorch, TensorFlow, and JAX.
ML 연구의 사실상 표준 프레임워크. dynamic graph·확장성·생태계 모두 1위.
Tensors and Dynamic neural networks in Python with strong GPU acceleration.
Apple Silicon용 MLX 런타임으로 Laya 타입 의사결정 모델을 초고속 실행한다. M3 Max에서 7-14ms 응답 속도를 내며, 텍스트 생성 없이 로컬에서 빠른 추론이 필요할 때 활용한다.
Native MLX runtime for Laya typed decision models — 7–14 ms short decisions on M3 Max. No text generation, PyTorch, or cloud API.
ML 모델에 빠르게 UI 붙이는 표준. HuggingFace Spaces 데모의 70% 이상이 이 프레임워크.
Build and share delightful machine learning apps — all in Python.
Rust로 만든 vector DB. 메모리 효율과 페이로드 필터링 성능이 뛰어남.
Qdrant — vector similarity search engine and database.
내 DB 스키마를 학습시켜 자연어 → 정확한 SQL을 만드는 라이브러리. 분석가 워크플로우에 강해요.
Chat with your SQL database — accurate Text-to-SQL Generation via LLMs using RAG.
Postgres에 벡터 검색 추가하는 extension. 별도 vector DB 없이 RAG 쉽게 시작.
Open-source vector similarity search for Postgres.
음성 인식의 표준. 99개 언어 지원하고 한국어 정확도가 매우 높아요.
Robust Speech Recognition via Large-Scale Weak Supervision.
OSS 모델 fine-tuning에서 가장 인기있는 도구. LoRA·QLoRA·Full FT 다 지원.
Go ahead and axolotl questions — fine-tuning toolkit.
Rust 기반의 lightweight 임베딩 DB. 멀티모달과 데이터 레이크 패턴에 강해요.
Developer-friendly, embedded retrieval engine for multimodal AI.
ML 학습·서빙·하이퍼파라미터 튜닝을 분산 처리하는 통합 프레임워크.
Unified framework for scaling AI and Python applications.
Airflow의 모던 대안. UX와 디버깅이 훨씬 좋고 데이터 사이언스 워크플로우에 친화적.
Modern workflow orchestration framework — easier than Airflow.
데이터 asset 중심 오케스트레이터. 데이터 lineage·observability에 강해요.
Dagster — a data orchestrator for the modern data stack.
RAG 시작용으로 가장 쉬운 vector DB. Python에서 5줄로 임베딩 저장·검색.
Chroma — the open-source embedding database.
스마트폰에서도 돌아가는 vision LLM. 작지만 GPT-4V 수준 작업도 가능.
MiniCPM-V — strong multimodal LLM for end-side deployment.
Netflix·Uber·Spotify 등이 공개한 프로덕션 ML 사례 모음. 실무 패턴 학습에 최고.
📚 Papers & tech blogs by companies sharing their work on data science & machine learning in production.
2.78조 파라미터 Kimi K3 모델을 단일 CPU와 8.24GB RAM으로 추론한다. BLAS, 프레임워크, GPU 없이 C99로만 구현되어 압도적인 휴대성과 경량 추론 환경을 제공한다.
A 2.78-trillion-parameter Kimi K3 running inference on a single CPU in 8.24 GB of RAM. Portable C99: no BLAS, no framework, no GPU.
분석가가 SQL로 데이터 변환을 작성하는 표준. modern data stack의 핵심.
dbt — Data Build Tool for analytics engineers.
Stable Diffusion LoRA 학습의 사실상 표준 도구. 캐릭터·스타일 학습이 직관적.
Training, generation and utility scripts for Stable Diffusion.
주간 업데이트되는 Python ML 라이브러리 랭킹. 새 도구 발견할 때 첫 출발점.
🏆 A ranked list of awesome machine learning Python libraries.
금융 도메인에 특화된 오픈소스 LLM 프로젝트. 뉴스 감성·시계열·로보어드바이저까지 노트북으로 다루고 있어요.
Open-Source Financial Large Language Models — democratizing internet-scale data for AI in finance.
전통적 ML 알고리즘의 표준 라이브러리. 분류·회귀·클러스터링 모두 한 곳에서.
Machine Learning in Python — classical ML algorithms.
GraphQL API 기반의 vector DB. 모듈식 ML 통합이 강점이고 RAG에 친화적.
Weaviate — open-source vector database that stores both objects and vectors.
알리바바의 HuggingFace 대안. 중국 모델·비디오 생성 모델이 풍부.
ModelScope — bring the notion of Model-as-a-Service to life.
LLM 벤치마크 표준. HellaSwag·MMLU 같은 평가를 한 번에 돌리는 프레임워크.
A framework for few-shot evaluation of language models.
구글의 ML 프레임워크. 프로덕션 deployment·TFLite·TFJS 같은 endpoint가 강점.
An Open Source Machine Learning Framework for Everyone.
Pandas/NumPy를 분산 처리로 확장하는 도구. 사이즈가 메모리를 넘을 때 첫 옵션.
Parallel computing with task scheduling.
Stability AI가 직접 푸시하는 모델 코드. 새 모델이 나오면 가장 먼저 여기에.
Generative Models by Stability AI — SDXL, SD3, Stable Cascade.
오픈 비전 LLM의 출발점. 이미지를 이해하는 LLM을 만들고 싶다면 첫 학습 자료.
Visual Instruction Tuning — LLaVA towards GPT-4V level.
Karpathy의 GPT 학습 코드. 단순함이 무기이고 LLM 내부를 직접 보고 싶을 때.
The simplest, fastest repository for training/finetuning medium-sized GPTs.
MS 리서치가 만든 다중 모델 오케스트레이터. HuggingFace 모델을 LLM이 도구로 호출.
JARVIS — connecting LLMs with ML community models.
import 한 줄만 바꾸면 Pandas 코드가 멀티 코어로 돌아가는 마법.
Modin — Scale your pandas workflows by changing one line of code.
CLIP 같은 비전-언어 모델을 적은 데이터로 파인튜닝하는 기법 구현체.
Conditional Prompt Learning for Vision-Language Models.
Salesforce가 만든 비전-언어 모델 모음. BLIP·CLIP 등이 통합 인터페이스로 묶여있어요.
LAVIS — a Library for Language-Vision Intelligence.
MS의 통합 음성-텍스트 모델. ASR·TTS·음성 변환 한 모델로.
Unified-Modal Speech-Text Pre-Training for Spoken Language Processing.
얼굴 인식을 위한 API입니다.
The world's simplest facial recognition api for Python and the command line
데이터를 분석하고 다루는 데 도움이 되는 파이썬 라이브러리입니다.
Flexible and powerful data analysis / manipulation library for Python, providing labeled data structures similar to R data.frame objects, statistical functions, and much more
데이터 시스템 사이의 zero-copy 데이터 교환 표준. 모든 빠른 dataframe의 베이스.
Apache Arrow is a multi-language toolbox for accelerated data interchange.
빅데이터 처리의 클래식. 여전히 페타바이트급에서 강력한 옵션.
Apache Spark - A unified analytics engine for large-scale data processing
Claude Mythos 아키텍처를 첫 번째 원칙부터 이론적으로 재구성한 오픈소스 프로젝트입니다.
A theoretical reconstruction of the Claude Mythos architecture, built from first principles using the available research literature.
OpenMMLab에서 개발한 객체 탐지 툴박스이자 벤치마크이다. 다양한 최신 객체 탐지 알고리즘을 구현하고 평가하는 데 사용된다.
OpenMMLab Detection Toolbox and Benchmark