본문으로 건너뛰기
AI 도감

코드 생성

자연어 설명에서 바로 실행되는 코드를 쓴다

코드와 소프트웨어입문 #33
입력텍스트코드

이 페이지의 본문은 영어로 제공됩니다. 제목과 요약은 한국어로 번역되었습니다.

이 능력이 뜻하는 것

Takes a description — write a function that reads a CSV into a list of dicts and skips blank lines — and outputs the corresponding source. Unlike code completion it targets a new feature or a whole piece of logic, usually written from scratch; unlike reasoning the deliverable is compilable, runnable code rather than a textual conclusion.

기술적으로 구현하는 방법

The base is a language model pre-trained on large code corpora; strict syntax makes code easier to verify by execution than prose, so unit-test pass signals work well as reward for reinforcement learning or rejection-sampling fine-tuning. Training and evaluation use problems with test cases such as HumanEval and MBPP, and generation can sample several candidates and discard the ones that fail.

대표 제품

11

GPT-4o

2024
OpenAI

네이티브 멀티모달 범용 모델. 텍스트·이미지·오디오를 한 창구에서 다룬다

모델 클로즈드 소스
텍스트이미지오디오텍스트오디오

Claude

2023
Anthropic

긴 문맥과 안전 정렬로 알려진 범용 대화 모델

모델 클로즈드 소스
텍스트이미지텍스트

DeepSeek-V3

2024
DeepSeek

총 671B, 토큰당 37B만 활성화하는 오픈웨이트 MoE

모델 공개 가중치
텍스트텍스트

Qwen

2023
Alibaba (Qwen)

여러 규모와 멀티모달 버전을 아우르는 오픈웨이트 모델 계열

모델 공개 가중치
텍스트이미지텍스트

Gemini

2023
Google DeepMind

네이티브 멀티모달에 초장문 문맥을 다루는 범용 모델

모델 클로즈드 소스
텍스트이미지오디오영상텍스트

GitHub Copilot

2021
Microsoft

편집기에서 문맥에 맞춰 코드를 이어 쓰고 고친다

도구 프리미엄
텍스트코드코드

o3

2025
OpenAI

답하기 전에 긴 추론을 거쳐, 추론 시 연산으로 정확도를 높인다

모델 클로즈드 소스
텍스트이미지텍스트

Claude Code

2025
Anthropic

터미널에서 여러 단계의 코딩 작업을 수행하는 에이전트

도구 클로즈드 소스
텍스트코드코드

DeepSeek-R1

2025
DeepSeek

강화학습으로 추론 사슬을 훈련하고 MIT 라이선스로 가중치를 공개한 추론 모델

모델 공개 가중치
텍스트텍스트

Devin

2024
Cognition (Devin)

셸·편집기·브라우저를 갖춘 자율 코딩 에이전트

앱 클로즈드 소스
텍스트코드코드

Cursor

2023
Anysphere (Cursor)

저장소 전체를 문맥으로 삼는 데스크톱 코드 편집기

앱 프리미엄
텍스트코드코드

관련 기관

대표적 용도

  • New features and utility scripts
  • Data cleaning and transformation scripts
  • Test cases and project scaffolding
  • Porting code across languages and frameworks

성능을 평가하는 방법

pass@k
Share of problems with at least one sample passing all tests out of k
Benchmark pass rate
Pass rate on fixed sets such as HumanEval
Compile and lint pass rate
Whether generated code passes the compiler and type checks as-is

경계와 난점

  • It invents libraries or methods, producing plausible interfaces that do not exist
  • Code can pass unit tests yet leave holes in edge cases, error handling and security
  • Multi-file changes are inconsistent, with call sites disagreeing with definitions

뒤에 있는 개념