MiniMax M3

3.00
Native multimodal, 1M context, MSA cuts cost 20x, built for multi-step complex tasks.
Advertisement 728 × 90
CompanyMinimax
Context1M
Released2026-05
Updated2026-07-20

MiniMax M3 Overview

MiniMax-M3 is a multimodal foundation model from MiniMax. It supports text, image, and video inputs with text output, a 1M-token context window, and is suited for long-horizon agentic work, coding, and tool use. It is built on MiniMax Sparse Attention (MSA), which replaces full attention with KV-block selection to cut per-token compute at long context — roughly 1/20 the cost of the previous generation at 1M tokens, with substantially faster prefill and decode while retaining quality across most tasks.

Trained as a native multimodal model on interleaved data and tuned for multi-turn, production-like collaboration via an interactive user-simulator framework, the model is oriented toward sustained, multi-step tasks rather than single-turn execution.

MiniMax M3 Pricing

PlanPriceDescription
Standard API Plan Input $0.60 / Output $2.40 per 1M Tokens Supports up to a 1M-token context window. Cached input is approximately $0.12 per 1M Tokens. Supports multimodal inputs including text, images, and video, making it suitable for developers and AI application integration.
Long Context API Plan Input $1.20 / Output $4.80 per 1M Tokens Applies when input context exceeds 512K Tokens. Designed for large-scale document analysis, complex reasoning tasks, and long-context applications.
Subscription Plans Approx. $20–$120/month Designed for individual users and developers, offering different usage limits and feature access depending on the plan.

MiniMax M3 Key Features

1. Top-Tier Coding & Autonomous Agent Capabilities

It can independently complete complex tasks like an AI engineer—for example, autonomously replicating AI research papers and optimizing code performance (achieving up to 10x acceleration in past tests). It has outperformed models like GPT-4.5 in relevant benchmarks.

2. 1M-Token Ultra-Long Context Window

It can process an entire book series like The Three-Body Problem trilogy or an entire code repository in one go, providing a solid foundation for long-horizon, multi-step complex tasks.

3. Native Multimodal & "Computer Use" Ability

It can understand image and video content, and like a human, it can "read" a computer screen and automatically operate software (such as ERP systems) to complete cross-application tasks.

Summary

1.Software engineers and developers. Equipped with top-tier coding capabilities, it can autonomously complete complex tasks such as replicating research papers and optimizing code. It outperforms GPT-4.5 in benchmark tests.

2.AI application and enterprise product teams. As the world's only open-source model that combines "coding + 1M context + multimodal" capabilities, it provides an ideal foundation for building highly autonomous Agents.

3.Enterprises sensitive to data privacy. Being open-source, it allows for private deployment, ensuring data security while delivering performance close to top closed-source models.

4.Cost-conscious developers and SMBs. The Token subscription plans offer excellent value (e.g., Max plan at 119 RMB/month includes 1.8 billion Tokens), making it suitable for high-frequency API calling scenarios.

Comments (0)

Leave a comment

Advertisement 728 × 90

Similar Models

Related News