한곳에서 제공업체별 최신 AI 모델을 살펴보세요. 각 호출은 가장 저렴하고 안정적인 옵션으로 지능형 라우팅되어 추가 작업 없이 더 좋은 가격을 얻습니다.
Seedance 2.5 text-to-video with 4-30s output, 480p/720p/1080p, synchronized audio, and optional web search.
Text-to-video generation with optional web search. 4-15s duration, 480p/720p/1080p, native audio sync.
Lightweight text-to-video generation. 4-15s duration, 480p/720p, native audio sync.
진짜 색 정확도, 구조화 작업, 분석적 비주얼 출력에 최적화된 OpenAI 고급 이미지 생성 모델. DALL-E 3 기술 기반.
Tongyi Wanxiang 3.0 video generation with text-to-video, image-to-video, and all-in-one reference-to-video. 480p/720p/1080p output, 2-30s or smart duration, per-second billing.
EvoLink 통합 API 라우트로 xAI 이미지 생성·편집을 이용하세요. 1–3장 참조 이미지, 1K/2K 출력, 13개 비율, 요청당 최대 10장을 지원합니다.
MiniMax H3 API. 텍스트로 영상 생성, 이미지로 영상 생성(첫/마지막 프레임), 레퍼런스로 영상 생성 세 라우트. 2K / 768p 출력, 4-15초, 출력 초당 과금.
Native multimodal model: text and image input with 1M context window. Same pricing as V4 Flash.
xAI의 최신 추론·도구 모델. 50만 Token 컨텍스트, Chat Completions와 Responses, 캐시 입력, 호출당 과금되는 5가지 서버 도구를 지원합니다.
Advanced reasoning model with thinking mode and 1M context window. Strongest DeepSeek tier.
Fast general-purpose model with 1M context window. Optional thinking mode.
리포지토리 규모의 코딩, 장기 실행 에이전트, 위험도 높은 리뷰를 위한 Anthropic의 최신 Opus 등급 플래그십. 100만 컨텍스트, 최대 출력 12.8만.
xAI Grok Imagine Video 1.5 Preview: text-to-video, image-to-video, and reference-to-video in one route — the number of input images (0, 1, or 2-7) selects the mode. 1-15s duration, 480p/720p/1080p.
Advanced Seedream 5.0 Pro image generation and editing with 1K/1.5K/2K output tiers, up to 10 reference images, and unified asynchronous task responses.
Tongyi Wanxiang image generation & editing 3.0 (DashScope). Text-to-image and image-to-image editing with up to 3 reference images, 1-6 outputs, flexible output sizes (auto, 1K/2K, aspect ratio, or custom pixels up to ~4.19MP), and smart prompt rewriting. Output billed per generated image with 1K and 2K priced the same; reference images billed separately.
Lite variant of Nano Banana 2 (Gemini 3.1 Flash-Lite Image): ~4s generation (2.7× faster than Nano Banana 2), native 1K output, and 14 aspect ratios. The fastest, lowest-cost tier for high-volume image generation and editing.
50만 Token 컨텍스트, Chat Completions, Responses, 캐시 입력 및 서버 도구를 지원하는 xAI 추론 모델입니다.
1,048,576 Token 컨텍스트와 Chat Completions 및 Messages API를 지원하는 Moonshot의 차세대 추론 모델입니다.
코딩과 에이전트를 위한 Google Flash 등급 주력 모델. 100만 Token 컨텍스트, 3.6 Flash보다 뛰어난 코드 생성과 터미널 실행, 프롬프트 캐싱, 텍스트·이미지·비디오·오디오 입력을 지원합니다.
Google's fast multimodal Flash model with a 1M context window, built-in reasoning, prompt caching, and unified text/image/video/audio input pricing.
Google's lowest-cost Flash Lite model with a 1M context window and multimodal input, built for high-volume, latency-sensitive, and batch workloads.
Gemini Omni Flash는 텍스트-비디오, 이미지-비디오, 레퍼런스-비디오, 비디오 편집을 하나의 비디오 API 엔드포인트로 제공하는 영상 생성 및 편집 모델입니다.
차세대 Nano Banana 모델로 더 세밀한 디테일과 뛰어난 프롬프트 준수로 우수한 이미지 생성을 제공합니다. 속도와 비용 효율을 최적화해 크리에이티브 전문가에게 큰 도약을 제공합니다.
Sol, Terra, Luna 세 등급으로 구성된 OpenAI 추론 제품군으로 1.05M 컨텍스트, 최대 128K 출력, 프롬프트 캐싱을 지원합니다.
Prompt-based AI audio generation for voice, dialogue, sound effects, music, and ambience, with optional reference audio guidance. Per-second billing.
코딩과 에이전트 작업을 위한 Anthropic의 가장 강력한 Sonnet. 1M 컨텍스트, 128K 최대 출력, 기본 활성화된 적응형 사고(adaptive thinking).
Midjourney V8.1 generates 4 images per request with HD/2K output options, reference-image workflows, and Draft/Fast speed tiers. Supports text-to-image and image-to-image through EvoLink.
입력 이미지가 무료인 가성비 Nano Banana 모델. 품질 대비 가격이 좋아 대량 이미지 생성에 적합합니다.
Premium text-to-image model with selectable 1K and 2K output. Built for crisp typography, cinematic lighting, and high-fidelity poster-grade compositions.
Kling 3.0 Turbo text-to-video — faster generation at 720P/1080P. Supports 3-15 second videos with per-second billing.
Z.ai's GLM-5.2 flagship text model for coding agents and agentic workflows, with a ~1M context window, deep thinking, tool calling, and prompt caching. Served over an OpenAI-compatible /v1/chat/completions endpoint.
Text-to-video generation. 3-15s duration, 720p/1080p, 9 aspect ratios, per-second billing.
Anthropic's most powerful and most intelligent Claude model — a new tier above Opus, with state-of-the-art reasoning, coding, and agentic capabilities. 1M context window.
Tongyi Wanxiang 2.7 video generation model with text-to-video, image-to-video, reference video, and video editing variants.
Text-to-video generation. 3-15s duration, 720p/1080p, per-second billing.
MiniMax's flagship multimodal text model with ~1M context, deep thinking, image/video/PDF input, and prompt caching. Available on both OpenAI-compatible (/v1/chat/completions) and Anthropic Messages (/v1/messages) endpoints.
Anthropic's most powerful Claude model with exceptional reasoning, coding, and agentic capabilities. 1M context window.
Google's next-gen Flash model with multimodal input (text/image/video/audio) at unified price and built-in reasoning output
알리바바의 플래그십 Max 모델로 100만 토큰 컨텍스트, 제어 가능한 사고, 프롬프트 캐싱을 3가지 API 프로토콜에서 제공합니다.
1M 컨텍스트 윈도우, 128K 최대 출력, Tool Search를 갖춘 OpenAI 최신 플래그십 모델.
Midjourney V7 generates 4 stunning images per request with native MJ prompt syntax. Supports text-to-image and image-to-image with three speed tiers (Draft/Fast/Turbo).
Multimodal content safety classifier for text and images. Detects 13 categories of harmful content (harassment, hate, sexual, violence, self-harm, illicit). OpenAI-compatible /v1/moderations endpoint.
AI-powered video super resolution. Enhance video quality with 1x, 2x, or 4x upscaling. Supports MP4 up to 50MB.
2K/3K 품질, 웹 검색 통합, 사용자 지정 픽셀 크기를 포함한 유연한 사이즈 옵션을 지원하는 차세대 AI 이미지 생성. 배치 생성(1~15장)을 지원합니다.
오디오 생성은 선택 사항. 텍스트-투-비디오 및 이미지-투-비디오, 4~12초, 480p/720p/1080p 품질.
Google's most cost-efficient model for high-volume agentic tasks, translation, classification, and data processing
Transfer human motion from a reference video onto a character in a reference image. Supports std/pro quality with per-second billing.
Kling O3 (V3 Omni) next-generation video model with text-to-video, image-to-video, reference-to-video, and video editing. Supports 3-15 second videos with per-second billing.
Kling 3.0 video model with text-to-video and image-to-video. Supports 3-15 second videos with per-second billing.
Tongyi Wanxiang 2.6 비디오 생성 모델로 텍스트-투-비디오, 이미지-투-비디오, 참조 비디오 변형을 지원합니다.
Google DeepMind 차세대 비디오 모델로 Fast와 Pro 변형을 제공합니다. 8초 비디오 생성과 향상된 품질을 지원합니다.
오디오 포함 OpenAI 최신 10~15초 비디오 생성 모델. 워터마크 제거 지원(가격 1.65배).
품질 대비 속도 비율이 뛰어난 초고속 이미지 생성 모델. 대량 생성에 적합하며 뛰어난 프롬프트 준수성을 제공합니다.
2K/4K 품질을 지원하는 고급 이미지 생성. 배치 생성(1~15장), 참조 이미지, 유연한 사이즈 옵션을 지원합니다.
텍스트-투-비디오 및 이미지-투-비디오를 지원하는 xAI Grok Imagine 비디오 생성 API. 6~30초 길이와 fun/normal/spicy 스타일 모드를 지원합니다.
보컬, 가사, 악기를 지원하는 AI 음악 생성. 여러 모델 버전으로 텍스트 프롬프트에서 프로급 곡을 생성합니다.
Google's advanced model optimized for custom tool calling and function execution with full reasoning capabilities
MiniMax's latest text model deployed on Alibaba Cloud, excelling at coding, office tasks, and text summarization with fast output speed. 204K context window with built-in reasoning capabilities.
1.05M 컨텍스트 윈도우, 128K 최대 출력, 고급 추론을 갖춘 코딩 및 에이전트 작업용 OpenAI 최신 플래그십 모델.
Google's latest iteration of Gemini 3 Pro with advanced multimodal capabilities and extended context support
코딩과 에이전트 작업을 위한 속도·지능·비용의 최적 균형. 200K 컨텍스트, 128K 최대 출력, Extended Thinking을 지원합니다.
BytePlus's latest LLM series with 256K context, tiered pricing by prompt length (32K/128K/256K), and cache billing. Available in Pro, Lite, Mini, and Code variants.
High-performance general-purpose chat model (DeepSeek-V3) with 128K context window and competitive pricing for everyday AI tasks
Advanced reasoning model (DeepSeek-R1) with chain-of-thought capabilities, 128K context window, optimized for complex problem-solving tasks
Google's most cost-efficient model for high-volume tasks like translation, classification, and data processing
Tongyi Wanxiang 이미지 생성 모델(Wan 2.5 Image)로 텍스트-투-이미지 및 이미지-투-이미지를 지원합니다.
Tongyi Wanxiang 비디오 생성 모델(Wan 2.5 Video)로 이미지-투-비디오와 텍스트-투-비디오를 지원합니다.
속도 3배의 Google 최상위 모델로, 오디오 입력 포함 멀티모달이며 비용 효율적입니다.
MiniMax Hailuo 2.3 API로 Fast/Standard 변형을 제공합니다. T2V/I2V에서 768p/1080p 출력을 지원합니다.
MiniMax Hailuo 02 - T2V, I2V, FLF 모드를 모두 지원. 512p/768p/1080p 해상도를 제공합니다.
400K 컨텍스트 윈도우, 128K 최대 출력, 고급 추론을 갖춘 코딩 및 에이전트 작업용 OpenAI 플래그십 모델.
OpenAI 최신 플래그십 모델로 고급 추론, 프롬프트 캐싱, 복잡한 작업을 위한 400K 컨텍스트 윈도우를 제공합니다.
Anthropic의 가장 강력한 Claude 모델로 뛰어난 추론, 코딩, 에이전트 기능과 200K 컨텍스트를 제공합니다.
Kling O1 비디오 생성 모델. 이미지-투-비디오, 비디오 편집, 빠른 비디오 편집 변형을 지원합니다. 참조 이미지로 스타일을 지시하며 3~20초 비디오를 생성합니다.
고급 멀티모달 기능과 확장된 컨텍스트를 갖춘 Google 차세대 언어 모델.
최적의 성능을 위해 200K 컨텍스트와 프롬프트 캐싱을 지원하는 빠르고 비용 효율적인 코딩 어시스턴트.
에이전트 구축과 코딩에 가장 지능적인 모델. 200K 컨텍스트, 확장된 사고, 고급 추론 기능을 제공합니다.
비용 최적화를 위한 프롬프트 캐싱 지원의 빠르고 효율적인 Google 언어 모델.
텍스트-투-비디오 및 이미지-투-비디오 기능의 고급 비디오 생성 모델. 2~12초, 720p/1080p 품질 옵션을 제공합니다.
확장 컨텍스트와 고급 추론 능력을 갖춘 Google 최강 언어 모델.
오디오 기반 립싱크를 지원하는 AI 디지털 휴먼 비디오 생성. 정지 이미지를 자연스러운 표정과 움직임의 말하는 아바타로 변환합니다.
Tongyi Qianwen 이미지 편집 모델로 지능형 이해와 다중 이미지 협업 편집을 지원합니다.
Gemini 2.5 Flash Image Preview는 자연어 기반 이미지 생성과 편집에 뛰어난 고급 AI 모델입니다.
4K 품질의 스토리 중심 이미지 생성: 다중 참조 융합과 실시간 편집으로 9개 이상의 일관된 비주얼을 생성합니다.
OpenAI named Astra as its next major model but has not linked it to GPT-6 or published an API. Track verified model IDs, pricing, limits, and EvoLink route readiness.
아직 발표되지 않은 Claude Fable 5.1 이름의 출시 현황을 추적합니다. 공식 제품명, API 이용, 요금, 안전장치와 EvoLink 경로 검증 상태를 확인하세요.
코딩과 장기 에이전트 작업을 위한 Z.ai의 최신 플래그십 모델입니다. 공식 glm-5.3 모델 ID, 확인된 사양, 검증된 EvoLink API 제공 상태를 확인하세요.
보도된 릴리스 상황, 검증된 EvoLink API 제공 여부, 모델 ID, 가격과 코딩 에이전트 평가 준비 상태를 확인하세요.
Gemini 3.5 Pro is testing with partners, but no public API route, model ID, pricing, or input specification is available yet.