WHAT IT IS
Step is the general model series StepFun has released since April 2023. StepFun was founded in Shanghai in 2023 by Jiang Daxin and others, with multimodality and on-device efficiency as main directions. Step-1 is a model in the hundred-billion-parameter range, and versions such as Step-1V accept image input; later releases add lightweight variants for phones and other devices. The series is offered through cloud interfaces and partly in open form, covering dialogue, image understanding and content generation.
Why it matters
It placed a hundred-billion-parameter general model and a lightweight version fit for phones on the same product line, betting on multimodality and on-device at once — an early, explicit version of that combination among start-up labs.
Key specs
- Parameters
- Hundred-billion range (Step-1)
- Modality
- Text, image in; text out
- Open weights
- Some versions open
- Released
- 2023-04
Capabilities
Related concepts
Transformer Architecture
Replacing word-by-word relay with a room where everyone speaks at once, so long-range dependencies are one hop away
Multimodal Generation
One model that learns to speak, to draw, to move — even to model the 3D world
Pretraining & Fine-tuning
Learn language first from vast unlabelled text, then specialise with little data — the most data-efficient paradigm in modern AI
Comparable products
Doubao
2023ByteDance’s general chat model and application
GLM
2023A Chinese general model that began with autoregressive blank-infilling pretraining
Kimi
2023A Chinese chat assistant known for long-context handling
MiniMax-M
2025An open-weight reasoning model with hybrid attention and a million-token context