プロバイダーを横断して今日のトップAIモデルを一か所で探索できます。各呼び出しは最も低コストで安定した選択肢へインテリジェントにルーティングされ、追加作業なしでより良い価格を得られます。
OpenAI の日常生成向けデフォルト:UI ドラフト、キャンペーン案、参考画像による修正。品質は max まで 5 段階、トークン課金。
OpenAI の精密編集向けモデル:商品参考画像、マスク、複数回の修正。品質は max まで 5 段階、トークン課金。
Seedance 2.5 text-to-video with 4-30s output, 480p/720p/1080p, synchronized audio, and optional web search.
EvoLink の統一 API で Nano Banana 2.1 の画像生成・編集を利用。1K・2K・4K、最大 14 枚の参照画像、任意のウェブ検索と画像検索に対応。
Text-to-video generation with optional web search. 4-15s duration, 480p/720p/1080p, native audio sync.
Lightweight text-to-video generation. 4-15s duration, 480p/720p, native audio sync.
真の色精度、構造化タスク、分析的ビジュアル出力に最適化されたOpenAIの高度な画像生成モデル。DALL-E 3技術ベース。
大量の分類・情報抽出、ルーティング、サブエージェントのタスク向けの Anthropic 最新の Haiku。プロンプトが 10 万トークン以下なら低い料金が適用されます。1M コンテキスト、128K 最大出力。
7種類のモードで既存動画を最大4Kに高画質化。出力動画の秒数に応じて課金。
複数ファイルのコード変更、文書への質問、ツール付きエージェント向けOpenAIモデル。共有コンテキスト1.05M、最大出力128K、プロンプトキャッシュに対応。ツールはResponses、Chat Completionsはツールなしで使用します。推論はlowから、maxはResponsesで利用します。
日常的なエージェント型コーディング、ツールを使うエージェント、高ボリュームの本番トラフィック向けの Anthropic 最新の Sonnet。公式価格は Sonnet 5 と同じ。1M コンテキスト、128K 最大出力。
Fast Seedream 5.0 Flash image generation and editing at one flat price per image across 1K/1.5K/2K, with up to 10 free reference images and unified asynchronous task responses.
Anthropic 最新の Opus モデル。長時間のエージェント型コーディング、ナレッジワーク、視覚分析に対応し、公式定価は Opus 5 より低い。100万コンテキスト、最大出力 12.8万。
複雑なコーディングとエージェントワークフロー向けの OpenAI GPT-6 モデル。1.05M コンテキスト、128K 最大出力、none から max までの推論強度、プロンプトキャッシュに対応。
OpenAI の GPT-6 で最も効率的なモデル。分類、抽出、ルーティングなど用途を絞った大量タスク向けで、1.05M コンテキスト、128K 最大出力、プロンプトキャッシュに対応。
xAI の最新の推論・ツール利用モデル。500,000 トークンのコンテキスト、Chat Completions と Responses、キャッシュ入力、従量課金のサーバーサイドツールに対応。
5.3 世代のスループット階層。テキスト・画像・動画・ファイル入力と 100 万トークンのコンテキストを、GLM-5.3 の約 4 分の 1 の価格で利用できます。
MiniMax H3 Max Turbo API for 5-15s text-to-video and first-frame animation with an optional last frame. Choose 480p drafts or 768p output with per-second pricing and a shared EvoLink balance.
Text and image input with a 1M context window and configurable thinking. Review current EvoLink pricing and protocol requirements before rollout.
OpenAI で最も高性能なモデル(GPT-6 Astra)。複雑な推論、エージェント型コーディング、コンピュータ操作、リサーチ向けで、1.05M コンテキスト、128K 最大出力、プロンプトキャッシュに対応。
コーディングとエージェント向けの Google 最高性能 Flash モデル。100 万 Token のコンテキスト、ソフトウェアエンジニアリング・エージェントタスク・多段推論で 3.7 Flash を大きく上回り、プロンプトキャッシュ、テキスト/画像/動画/音声入力に対応。
Official MiniMax H3 Max video generation through EvoLink, with text-to-video, first-frame image-to-video and image/video/audio reference routes. 768p / 480p, 5-15s; reference inputs have separate fees.
Anthropic 最高性能のモデル(Opus の上位となる Fable 帯)。リポジトリ規模のコーディング、長時間エージェント、重要度の高いレビューに対応。100万コンテキスト、最大出力 12.8万。
Gemini Omni 1.1 Flash text-to-video with 3-10s output, 360p/720p/1080p/4K, and natively synchronized audio.
Tongyi Wanxiang 3.0 video generation with text-to-video, image-to-video, and all-in-one reference-to-video. 480p/720p/1080p output, 2-30s or smart duration, per-second billing.
Tongyi Wanxiang 3.0 Prime is the speed-oriented tier for significantly faster end-to-end video generation. Use three EvoLink model IDs for text, image, or all-in-one reference workflows up to 30 seconds.
EvoLink の統一 API ルートで xAI の画像生成と編集を利用。1〜3 枚の参照画像、1K/2K 出力、13 種類の比率、1 リクエスト最大 10 枚に対応します。
MiniMax H3 API。テキストから動画、画像から動画(開始/終了フレーム)、参照素材から動画の3ルート。出力は2K/768p・4〜15秒、出力秒単位の課金です。
コーディングエージェントと長期エンジニアリング向けの Z.ai フラッグシップ推論モデル。100 万トークンのコンテキスト、3 段階の常時オン思考、プロンプトキャッシュ、Chat Completions・Responses・Anthropic Messages の各エンドポイントに対応。
5.3 世代のネイティブマルチモーダル大量処理階層。テキスト・画像・動画・ファイル入力に対応し、100 万トークンのコンテキストを GLM-5.3 の約 10 分の 1 の価格で。
Legacy Vision Exp ID, now routed to DeepSeek V4.1 Flash for text and image input.
xAI の最新推論モデル。50万 Token のコンテキスト、Chat Completions と Responses、キャッシュ入力、従量課金の 5 種類のサーバーツールに対応。
Advanced reasoning model with thinking mode and 1M context window. Strongest DeepSeek tier.
Fast general-purpose model with 1M context window. Optional thinking mode.
Anthropic の前世代 Opus 系フラッグシップ。リポジトリ規模のコーディング、長時間エージェント、重要度の高いレビューに対応。100万コンテキスト、最大出力 12.8万。
xAI Grok Imagine Video 1.5 Preview: text-to-video, image-to-video, and reference-to-video in one route — the number of input images (0, 1, or 2-7) selects the mode. 1-15s duration, 480p/720p/1080p.
Advanced Seedream 5.0 Pro image generation and editing with 1K/1.5K/2K output tiers, up to 10 reference images, and unified asynchronous task responses.
Tongyi Wanxiang image generation & editing 3.0 (DashScope). Text-to-image and image-to-image editing with up to 3 reference images, 1-6 outputs, flexible output sizes (auto, 1K/2K, aspect ratio, or custom pixels up to ~4.19MP), and smart prompt rewriting. Output billed per generated image with 1K and 2K priced the same; reference images billed separately.
Lite variant of Nano Banana 2 (Gemini 3.1 Flash-Lite Image): ~4s generation (2.7× faster than Nano Banana 2), native 1K output, and 14 aspect ratios. The fastest, lowest-cost tier for high-volume image generation and editing.
50万 Token のコンテキスト、Chat Completions、Responses、キャッシュ入力、サーバーツールに対応する xAI の推論モデル。
1,048,576 Token のコンテキストと Chat Completions・Responses・Messages の各 API に対応する Moonshot の次世代推論モデル。
コーディングとエージェント向けの Google Flash クラス主力モデル。100 万 Token のコンテキスト、3.6 Flash を上回るコード生成とターミナル実行、プロンプトキャッシュ、テキスト/画像/動画/音声入力に対応。
Google's fast multimodal Flash model with a 1M context window, built-in reasoning, prompt caching, and unified text/image/video/audio input pricing.
Google's lowest-cost Flash Lite model with a 1M context window and multimodal input, built for high-volume, latency-sensitive, and batch workloads.
Gemini Omni Flash は、テキスト動画生成、画像動画生成、参照画像動画生成、動画編集を単一の動画 API エンドポイントで扱える動画生成・編集モデルです。
次世代Nano Bananaモデルで、より精細で高いプロンプト追従性を備えた優れた画像生成を提供。速度とコスト効率を両立し、クリエイティブプロにとって大きな飛躍をもたらします。
Sol、Terra、Luna の 3 モデルを備えた OpenAI 推論ファミリー。1.05M コンテキスト、最大 128K 出力、プロンプトキャッシュに対応。
Prompt-based AI audio generation for voice, dialogue, sound effects, music, and ambience, with optional reference audio guidance. Per-second billing.
コーディングとエージェントタスク向けの Anthropic の前世代 Sonnet。1M コンテキスト、128K 最大出力、adaptive thinking をデフォルトで有効化。
EvoLink の 1 つの API ルートで Midjourney V8.2 を利用:1 リクエスト 4 枚、Draft/Fast 速度、HD 出力、無料の --q ディテールパラメータ、さらにバリエーション・リミックス・キャンバス編集・リテクスチャ・アップロードペイント・背景除去のモデル ID。
Midjourney V8.1 generates 4 images per request with HD/2K output options, reference-image workflows, and Draft/Fast speed tiers. Supports text-to-image and image-to-image through EvoLink.
入力画像が無料のコスト効率に優れたNano Bananaモデル。品質/価格比が高く大量生成に最適です。
Premium text-to-image model with selectable 1K and 2K output. Built for crisp typography, cinematic lighting, and high-fidelity poster-grade compositions.
Kling 3.0 Turbo text-to-video — faster generation at 720P/1080P. Supports 3-15 second videos with per-second billing.
Z.ai's GLM-5.2 flagship text model for coding agents and agentic workflows, with a ~1M context window, deep thinking, tool calling, and prompt caching. Served over OpenAI-compatible /v1/chat/completions and /v1/responses endpoints.
Text-to-video generation. 3-15s duration, 720p/1080p, 9 aspect ratios, per-second billing.
Anthropic's previous-generation Fable model above the Opus tier, built for demanding reasoning, coding, and agentic workloads. 1M context window.
Tongyi Wanxiang 2.7 video generation model with text-to-video, image-to-video, reference video, and video editing variants.
Text-to-video generation. 3-15s duration, 720p/1080p, per-second billing.
MiniMax's flagship multimodal text model with ~1M context, deep thinking, image/video/PDF input, and prompt caching. Available on both OpenAI-compatible (/v1/chat/completions) and Anthropic Messages (/v1/messages) endpoints.
Anthropic's most powerful Claude model with exceptional reasoning, coding, and agentic capabilities. 1M context window.
Google's next-gen Flash model with multimodal input (text/image/video/audio) at unified price and built-in reasoning output
アリババのフラッグシップMaxモデル。100万トークンのコンテキスト、制御可能な思考、プロンプトキャッシュを3つのAPIプロトコルで提供します。
1Mコンテキスト、128K最大出力、Tool Search を備えたOpenAI最新フラッグシップモデル。
Midjourney V7 generates 4 stunning images per request with native MJ prompt syntax. Supports text-to-image and image-to-image with three speed tiers (Draft/Fast/Turbo).
Multimodal content safety classifier for text and images. Detects 13 categories of harmful content (harassment, hate, sexual, violence, self-harm, illicit). OpenAI-compatible /v1/moderations endpoint.
2K/3K品質、ウェブ検索統合、カスタムピクセル寸法を含む柔軟なサイズオプションに対応した次世代AI画像生成。バッチ生成(1〜15枚)に対応。
オーディオ生成は任意。テキスト→動画/画像→動画に対応。4〜12秒、480p/720p/1080p品質。
Google's most cost-efficient model for high-volume agentic tasks, translation, classification, and data processing
Transfer human motion from a reference video onto a character in a reference image. Supports std/pro quality with per-second billing.
Transfer reference-video motion onto a character image with Kling 2.6 Motion Control. 720p/1080p output; 3–10s or 3–30s references by orientation; actual output may be shorter. Per-second billing.
Kling O3 (V3 Omni) next-generation video model with text-to-video, image-to-video, reference-to-video, and video editing. Supports 3-15 second videos with per-second billing.
Kling 3.0 video model with text-to-video and image-to-video. Supports 3-15 second videos with per-second billing.
通義万相2.6の動画生成モデル。テキスト→動画、画像→動画、参照動画バリエーションに対応。
Google DeepMind次世代動画モデル。FastとProバリアントを備え、8秒動画生成と品質向上に対応。
音声付きのOpenAI最新10〜15秒動画生成モデル。ウォーターマーク削除に対応(価格1.65倍)。
品質と速度の比率が優れた超高速画像生成モデル。高ボリューム生成に最適で高いプロンプト追従性を備えます。
2K/4K品質に対応した高度な画像生成。バッチ生成(1〜15枚)、参照画像、柔軟なサイズオプションに対応。
テキストから動画・画像から動画に対応するxAI Grok Imagine動画生成API。6〜30秒の長さとfun/normal/spicyのスタイルモードに対応。
ボーカル、歌詞、伴奏対応のAI音楽生成。複数モデル版でテキストからプロ品質の楽曲を作成できます。
Google's advanced model optimized for custom tool calling and function execution with full reasoning capabilities
MiniMax's latest text model deployed on Alibaba Cloud, excelling at coding, office tasks, and text summarization with fast output speed. 204K context window with built-in reasoning capabilities.
1.05Mコンテキスト、128K最大出力、先進推論を備えたコーディングおよびエージェントタスク向けOpenAI最新フラッグシップモデル。
Google's latest iteration of Gemini 3 Pro with advanced multimodal capabilities and extended context support
コーディングとエージェントタスクにおける速度・知能・コストの最適なバランス。200Kコンテキスト、128K最大出力、Extended Thinkingに対応。
BytePlus's latest LLM series with 256K context, tiered pricing by prompt length (32K/128K/256K), and cache billing. Available in Pro, Lite, Mini, and Code variants.
Legacy DeepSeek Chat ID, now routed to DeepSeek V4.1 Flash. Use deepseek-v4.1-flash for new integrations.
Legacy DeepSeek Reasoner ID, now routed to DeepSeek V4.1 Flash. Use deepseek-v4.1-flash for new integrations.
Google's most cost-efficient model for high-volume tasks like translation, classification, and data processing
通義万相の画像生成モデル(Wan 2.5 Image)。テキスト→画像と画像→画像に対応。
通義万相の動画生成モデル(Wan 2.5 Video)。画像→動画とテキスト→動画に対応。
3倍速度のGoogle最速フロンティアモデル。音声入力を含むマルチモーダルで、コスト効率も高い。
MiniMax Hailuo 2.3 API。Fast/Standardバリアント、T2V/I2Vで768p/1080p出力に対応。
MiniMax Hailuo 02 - T2V、I2V、FLFのフル機能。512p/768p/1080p解像度に対応。
400Kコンテキスト、128K最大出力、先進推論を備えたコーディングおよびエージェントタスク向けOpenAIフラッグシップモデル。
OpenAI最新フラッグシップモデル。高度推論、プロンプトキャッシュ、複雑タスク向け400Kコンテキスト。
Anthropic最強のClaudeモデル。卓越した推論、コーディング、エージェント機能を備え、200Kコンテキストに対応。
Kling O1の動画生成モデル。画像→動画、動画編集、高速動画編集に対応。参照画像によるスタイル指示で3〜20秒の動画を生成します。
先進的なマルチモーダル機能と拡張コンテキストを備えたGoogle次世代言語モデル。
最適なパフォーマンスのため200Kコンテキストとプロンプトキャッシュに対応した高速・低コストのコーディングアシスタント。
エージェント構築とコーディングに最も賢いモデル。200Kコンテキスト、拡張思考、高度推論を搭載。
コスト最適化のためのプロンプトキャッシュに対応した高速・効率的なGoogle言語モデル。
テキスト→動画/画像→動画に対応する高度な動画生成モデル。2〜12秒、720p/1080p品質オプション。
拡張コンテキストと高度推論を備えたGoogle最強の言語モデル。
音声駆動リップシンクを備えたAIデジタルヒューマン動画生成。静止画を自然な表情と動きの話すアバターに変換します。
通義千問の画像編集モデル。高度な理解と複数画像の協調編集に対応。
Gemini 2.5 Flash Image Previewは自然言語駆動の画像生成・編集に優れた先進AIモデルです。
4K品質のストーリー駆動型画像生成。複数参照の融合とリアルタイム編集で、9枚以上の一貫したビジュアルを作成します。
EvoLinkで連携準備中。コーディング・調査・ファイル処理向けのマネージドAgentサービス。提供開始の通知を受け取れます。
報道されたリリース状況、検証済みのEvoLink API提供状況、モデルID、価格、コーディングエージェント評価の準備状況を確認できます。
EvoLinkからのアクセスは未検証です。本番ルートを変更する前に、公開の根拠と連携要件を確認してください。
Google announced Gemini 4 Argon with restricted access, a 1M-token output limit and introductory token rates. Public API and EvoLink access remain pending verification.
Gemini 3.5 Pro is testing with partners, but no public API route, model ID, pricing, or input specification is available yet.
Plan Kling 4.0 Flash API integration. EvoLink access, API pricing and request identifiers are not verified; app access is separate.