プロバイダーを横断して今日のトップAIモデルを一か所で探索できます。各呼び出しは最も低コストで安定した選択肢へインテリジェントにルーティングされ、追加作業なしでより良い価格を得られます。
Seedance 2.5 text-to-video with 4-30s output, 480p/720p/1080p, synchronized audio, and optional web search.
Text-to-video generation with optional web search. 4-15s duration, 480p/720p/1080p, native audio sync.
Lightweight text-to-video generation. 4-15s duration, 480p/720p, native audio sync.
真の色精度、構造化タスク、分析的ビジュアル出力に最適化されたOpenAIの高度な画像生成モデル。DALL-E 3技術ベース。
Tongyi Wanxiang 3.0 video generation with text-to-video, image-to-video, and all-in-one reference-to-video. 480p/720p/1080p output, 2-30s or smart duration, per-second billing.
EvoLink の統一 API ルートで xAI の画像生成と編集を利用。1〜3 枚の参照画像、1K/2K 出力、13 種類の比率、1 リクエスト最大 10 枚に対応します。
MiniMax H3 API。テキストから動画、画像から動画(開始/終了フレーム)、参照素材から動画の3ルート。出力は2K/768p・4〜15秒、出力秒単位の課金です。
Native multimodal model: text and image input with 1M context window. Same pricing as V4 Flash.
xAI の最新推論モデル。50万 Token のコンテキスト、Chat Completions と Responses、キャッシュ入力、従量課金の 5 種類のサーバーツールに対応。
Advanced reasoning model with thinking mode and 1M context window. Strongest DeepSeek tier.
Fast general-purpose model with 1M context window. Optional thinking mode.
Anthropic 最新の Opus 系フラッグシップ。リポジトリ規模のコーディング、長時間エージェント、重要度の高いレビューに対応。100万コンテキスト、最大出力 12.8万。
xAI Grok Imagine Video 1.5 Preview: text-to-video, image-to-video, and reference-to-video in one route — the number of input images (0, 1, or 2-7) selects the mode. 1-15s duration, 480p/720p/1080p.
Advanced Seedream 5.0 Pro image generation and editing with 1K/1.5K/2K output tiers, up to 10 reference images, and unified asynchronous task responses.
Tongyi Wanxiang image generation & editing 3.0 (DashScope). Text-to-image and image-to-image editing with up to 3 reference images, 1-6 outputs, flexible output sizes (auto, 1K/2K, aspect ratio, or custom pixels up to ~4.19MP), and smart prompt rewriting. Output billed per generated image with 1K and 2K priced the same; reference images billed separately.
Lite variant of Nano Banana 2 (Gemini 3.1 Flash-Lite Image): ~4s generation (2.7× faster than Nano Banana 2), native 1K output, and 14 aspect ratios. The fastest, lowest-cost tier for high-volume image generation and editing.
50万 Token のコンテキスト、Chat Completions、Responses、キャッシュ入力、サーバーツールに対応する xAI の推論モデル。
1,048,576 Token のコンテキストと Chat Completions・Messages の両 API に対応する Moonshot の次世代推論モデル。
コーディングとエージェント向けの Google Flash クラス主力モデル。100 万 Token のコンテキスト、3.6 Flash を上回るコード生成とターミナル実行、プロンプトキャッシュ、テキスト/画像/動画/音声入力に対応。
Google's fast multimodal Flash model with a 1M context window, built-in reasoning, prompt caching, and unified text/image/video/audio input pricing.
Google's lowest-cost Flash Lite model with a 1M context window and multimodal input, built for high-volume, latency-sensitive, and batch workloads.
Gemini Omni Flash は、テキスト動画生成、画像動画生成、参照画像動画生成、動画編集を単一の動画 API エンドポイントで扱える動画生成・編集モデルです。
次世代Nano Bananaモデルで、より精細で高いプロンプト追従性を備えた優れた画像生成を提供。速度とコスト効率を両立し、クリエイティブプロにとって大きな飛躍をもたらします。
Sol、Terra、Luna の 3 モデルを備えた OpenAI 推論ファミリー。1.05M コンテキスト、最大 128K 出力、プロンプトキャッシュに対応。
Prompt-based AI audio generation for voice, dialogue, sound effects, music, and ambience, with optional reference audio guidance. Per-second billing.
コーディングとエージェントタスク向けの Anthropic 最も高性能な Sonnet。1M コンテキスト、128K 最大出力、adaptive thinking をデフォルトで有効化。
Midjourney V8.1 generates 4 images per request with HD/2K output options, reference-image workflows, and Draft/Fast speed tiers. Supports text-to-image and image-to-image through EvoLink.
入力画像が無料のコスト効率に優れたNano Bananaモデル。品質/価格比が高く大量生成に最適です。
Premium text-to-image model with selectable 1K and 2K output. Built for crisp typography, cinematic lighting, and high-fidelity poster-grade compositions.
Kling 3.0 Turbo text-to-video — faster generation at 720P/1080P. Supports 3-15 second videos with per-second billing.
Z.ai's GLM-5.2 flagship text model for coding agents and agentic workflows, with a ~1M context window, deep thinking, tool calling, and prompt caching. Served over an OpenAI-compatible /v1/chat/completions endpoint.
Text-to-video generation. 3-15s duration, 720p/1080p, 9 aspect ratios, per-second billing.
Anthropic's most powerful and most intelligent Claude model — a new tier above Opus, with state-of-the-art reasoning, coding, and agentic capabilities. 1M context window.
Tongyi Wanxiang 2.7 video generation model with text-to-video, image-to-video, reference video, and video editing variants.
Text-to-video generation. 3-15s duration, 720p/1080p, per-second billing.
MiniMax's flagship multimodal text model with ~1M context, deep thinking, image/video/PDF input, and prompt caching. Available on both OpenAI-compatible (/v1/chat/completions) and Anthropic Messages (/v1/messages) endpoints.
Anthropic's most powerful Claude model with exceptional reasoning, coding, and agentic capabilities. 1M context window.
Google's next-gen Flash model with multimodal input (text/image/video/audio) at unified price and built-in reasoning output
アリババのフラッグシップMaxモデル。100万トークンのコンテキスト、制御可能な思考、プロンプトキャッシュを3つのAPIプロトコルで提供します。
1Mコンテキスト、128K最大出力、Tool Search を備えたOpenAI最新フラッグシップモデル。
Midjourney V7 generates 4 stunning images per request with native MJ prompt syntax. Supports text-to-image and image-to-image with three speed tiers (Draft/Fast/Turbo).
Multimodal content safety classifier for text and images. Detects 13 categories of harmful content (harassment, hate, sexual, violence, self-harm, illicit). OpenAI-compatible /v1/moderations endpoint.
AI-powered video super resolution. Enhance video quality with 1x, 2x, or 4x upscaling. Supports MP4 up to 50MB.
2K/3K品質、ウェブ検索統合、カスタムピクセル寸法を含む柔軟なサイズオプションに対応した次世代AI画像生成。バッチ生成(1〜15枚)に対応。
オーディオ生成は任意。テキスト→動画/画像→動画に対応。4〜12秒、480p/720p/1080p品質。
Google's most cost-efficient model for high-volume agentic tasks, translation, classification, and data processing
Transfer human motion from a reference video onto a character in a reference image. Supports std/pro quality with per-second billing.
Kling O3 (V3 Omni) next-generation video model with text-to-video, image-to-video, reference-to-video, and video editing. Supports 3-15 second videos with per-second billing.
Kling 3.0 video model with text-to-video and image-to-video. Supports 3-15 second videos with per-second billing.
通義万相2.6の動画生成モデル。テキスト→動画、画像→動画、参照動画バリエーションに対応。
Google DeepMind次世代動画モデル。FastとProバリアントを備え、8秒動画生成と品質向上に対応。
音声付きのOpenAI最新10〜15秒動画生成モデル。ウォーターマーク削除に対応(価格1.65倍)。
品質と速度の比率が優れた超高速画像生成モデル。高ボリューム生成に最適で高いプロンプト追従性を備えます。
2K/4K品質に対応した高度な画像生成。バッチ生成(1〜15枚)、参照画像、柔軟なサイズオプションに対応。
テキストから動画・画像から動画に対応するxAI Grok Imagine動画生成API。6〜30秒の長さとfun/normal/spicyのスタイルモードに対応。
ボーカル、歌詞、伴奏対応のAI音楽生成。複数モデル版でテキストからプロ品質の楽曲を作成できます。
Google's advanced model optimized for custom tool calling and function execution with full reasoning capabilities
MiniMax's latest text model deployed on Alibaba Cloud, excelling at coding, office tasks, and text summarization with fast output speed. 204K context window with built-in reasoning capabilities.
1.05Mコンテキスト、128K最大出力、先進推論を備えたコーディングおよびエージェントタスク向けOpenAI最新フラッグシップモデル。
Google's latest iteration of Gemini 3 Pro with advanced multimodal capabilities and extended context support
コーディングとエージェントタスクにおける速度・知能・コストの最適なバランス。200Kコンテキスト、128K最大出力、Extended Thinkingに対応。
BytePlus's latest LLM series with 256K context, tiered pricing by prompt length (32K/128K/256K), and cache billing. Available in Pro, Lite, Mini, and Code variants.
High-performance general-purpose chat model (DeepSeek-V3) with 128K context window and competitive pricing for everyday AI tasks
Advanced reasoning model (DeepSeek-R1) with chain-of-thought capabilities, 128K context window, optimized for complex problem-solving tasks
Google's most cost-efficient model for high-volume tasks like translation, classification, and data processing
通義万相の画像生成モデル(Wan 2.5 Image)。テキスト→画像と画像→画像に対応。
通義万相の動画生成モデル(Wan 2.5 Video)。画像→動画とテキスト→動画に対応。
3倍速度のGoogle最速フロンティアモデル。音声入力を含むマルチモーダルで、コスト効率も高い。
MiniMax Hailuo 2.3 API。Fast/Standardバリアント、T2V/I2Vで768p/1080p出力に対応。
MiniMax Hailuo 02 - T2V、I2V、FLFのフル機能。512p/768p/1080p解像度に対応。
400Kコンテキスト、128K最大出力、先進推論を備えたコーディングおよびエージェントタスク向けOpenAIフラッグシップモデル。
OpenAI最新フラッグシップモデル。高度推論、プロンプトキャッシュ、複雑タスク向け400Kコンテキスト。
Anthropic最強のClaudeモデル。卓越した推論、コーディング、エージェント機能を備え、200Kコンテキストに対応。
Kling O1の動画生成モデル。画像→動画、動画編集、高速動画編集に対応。参照画像によるスタイル指示で3〜20秒の動画を生成します。
先進的なマルチモーダル機能と拡張コンテキストを備えたGoogle次世代言語モデル。
最適なパフォーマンスのため200Kコンテキストとプロンプトキャッシュに対応した高速・低コストのコーディングアシスタント。
エージェント構築とコーディングに最も賢いモデル。200Kコンテキスト、拡張思考、高度推論を搭載。
コスト最適化のためのプロンプトキャッシュに対応した高速・効率的なGoogle言語モデル。
テキスト→動画/画像→動画に対応する高度な動画生成モデル。2〜12秒、720p/1080p品質オプション。
拡張コンテキストと高度推論を備えたGoogle最強の言語モデル。
音声駆動リップシンクを備えたAIデジタルヒューマン動画生成。静止画を自然な表情と動きの話すアバターに変換します。
通義千問の画像編集モデル。高度な理解と複数画像の協調編集に対応。
Gemini 2.5 Flash Image Previewは自然言語駆動の画像生成・編集に優れた先進AIモデルです。
4K品質のストーリー駆動型画像生成。複数参照の融合とリアルタイム編集で、9枚以上の一貫したビジュアルを作成します。
OpenAI named Astra as its next major model but has not linked it to GPT-6 or published an API. Track verified model IDs, pricing, limits, and EvoLink route readiness.
未発表の Claude Fable 5.1 を追跡するページです。正式名称、API 提供、料金、安全機能、EvoLink ルートの検証状況を確認できます。
コーディングと長期エージェントタスク向けの Z.ai 最新フラッグシップモデル。公式モデル ID glm-5.3、確認済みの仕様、検証済みの EvoLink API 提供状況を確認できます。
報道されたリリース状況、検証済みのEvoLink API提供状況、モデルID、価格、コーディングエージェント評価の準備状況を確認できます。
Gemini 3.5 Pro is testing with partners, but no public API route, model ID, pricing, or input specification is available yet.