추론과 사고 연쇄
어려운 문제를 중간 단계로 나눠 푼다
이 페이지의 본문은 영어로 제공됩니다. 제목과 요약은 한국어로 번역되었습니다.
이 능력이 뜻하는 것
Takes a problem requiring several steps — a maths problem, a logic puzzle, a task needing a plan — and returns the answer together with intermediate steps. Unlike free generation the goal is correctness rather than fluency, and unlike ordinary conversation it spends much of its compute on an internal scratchpad, showing the user mainly the result and a short rationale.
기술적으로 구현하는 방법
One family relies on prompting: examples or instructions ask the model to think before answering, writing steps out explicitly — chain-of-thought. Another relies on training: large-scale reinforcement learning on verifiable problems such as maths and code teaches the model to generate a longer deliberation before committing. Self-consistency voting over several sampled chains, and combining deliberation with tool calls or search, are now common too.
대표 제품
23o3
2025답하기 전에 긴 추론을 거쳐, 추론 시 연산으로 정확도를 높인다
DeepSeek-R1
2025강화학습으로 추론 사슬을 훈련하고 MIT 라이선스로 가중치를 공개한 추론 모델
Gemini
2023네이티브 멀티모달에 초장문 문맥을 다루는 범용 모델
Claude
2023긴 문맥과 안전 정렬로 알려진 범용 대화 모델
Qwen
2023여러 규모와 멀티모달 버전을 아우르는 오픈웨이트 모델 계열
MiniMax-M
2025하이브리드 어텐션과 백만 토큰 문맥을 갖춘 오픈웨이트 추론 모델
DeepSeek-V3
2024총 671B, 토큰당 37B만 활성화하는 오픈웨이트 MoE
GPT-4o
2024네이티브 멀티모달 범용 모델. 텍스트·이미지·오디오를 한 창구에서 다룬다
Jamba
2024상태공간 모델과 Transformer를 혼합한 오픈웨이트 모델
DBRX
2024Databricks의 오픈 웨이트 MoE 언어 모델
Mistral Large
2024유럽 오픈웨이트 연구소의 플래그십 상용 모델
Gemini
2024검색·오피스·멀티모달 모델을 하나의 대화 창구로 모았다
Grok
2023소셜 플랫폼 데이터와 결합된 대화 모델
Yi
2023중영 이중언어, 초장문 문맥 버전을 제공하는 오픈웨이트 모델
Hunyuan
2023오픈웨이트 버전을 포함한 텐센트의 범용 모델 계열
Doubao
2023바이트댄스의 범용 대화 모델과 앱
Phi
2023작은 규모와 선별 데이터로 만든 효율적 소형 모델
Step
2023멀티모달과 온디바이스 지향 범용 모델 계열
Baichuan
2023중국어 환경을 겨냥한 오픈웨이트 범용 모델
GLM
2023자기회귀 빈칸 채우기 사전학습에서 출발한 중국어 범용 모델
Llama
2023개방형 가중치 노선을 주류로 만든 범용 모델 계열
ChatGPT
2022대규모 언어 모델을 누구나 쓰는 대화 창으로 만들었다
Pangu
2020화웨이의 판구 기반 모델 시리즈
관련 기관
대표적 용도
- Solving maths and physics problems
- Multi-step planning and scheduling
- Code and data problems needing derivation
- Reasoned choices under constraints
성능을 평가하는 방법
- Math benchmark accuracy
- Correct-answer rate on sets such as GSM8K, MATH, AIME
- pass@k
- Share of problems solved by at least one of k samples
- Chain-of-thought faithfulness
- Whether the written steps actually reflect how the answer was reached
경계와 난점
- The written reasoning can be unfaithful: the answer comes first and a plausible rationale is added after
- In long chains an early misstep propagates, yet the final answer still sounds assured
- Deliberating at length on trivial questions inflates latency and cost