Skip to main content

Available Models List

CloudBase AI comes with built-in large language models from major providers. You can call them directly through the Mini Program, Web SDK, Node SDK, HTTP API, and more (see Invocation Methods) — there is no need to integrate each provider yourself.

The built-in models' capabilities are aligned with each vendor's official models: CloudBase standardizes the invocation path, multi-platform SDKs, and protocols, but does not trim, replace, or alter the models' own capabilities. The capabilities, parameters, and limits described in each vendor's official documentation therefore apply on CloudBase as well.

The models actually available and their prices are subject to what is shown in the CloudBase console. For models that are no longer served, see Deprecated Models List.

How to read the tables
  • ✅ means the vendor's official documentation explicitly states support for that capability; ❌ means the official documentation explicitly states it is not supported; — means the vendor does not state it for that specific model, which does not necessarily mean it is unsupported.
  • "Context" is the model's maximum context window, provided as a selection reference.
  • Each vendor section links to its official documentation, which is the source of the capability values. If a value differs from the vendor's latest documentation or the CloudBase console, the latter prevails.

Models and Capabilities​

Hunyuan​

Official documentation: Hunyuan invocation guide

ModelContextImageVideoFileDeep ThinkingTool CallingStructured OutputCaching
hy4-preview1M❌❌—✅✅✅✅
hy3256K❌❌—✅✅✅✅
hy-vision-2.0-instruct44K✅❌—————
hunyuan-t1-vision-2025091640K✅❌—✅———
hunyuan-turbos-vision-video-2025072832K❌✅—————
hunyuan-role-latest32K———————
hy-role32K———————
hy-mt2-lite8K———————
hy-mt2-plus8K———————
hy-mt2-pro8K———————

Of these, hunyuan-role-latest and hy-role target role-playing, while hy-mt2-lite, hy-mt2-plus, and hy-mt2-pro target translation. They are tuned for specific tasks, so the capability columns above do not apply.

  • Model specs: The TokenHub model list gives each model's capability support, context window, and max input / output
  • Deep thinking: hy4-preview and hy3 have deep thinking on by default and support Preserved Thinking; toggle it with thinking.type, and set reasoning_effort to low / high (default high). With tools, hy3 adapts to task complexity and maps low to high
  • Structured output: Emit JSON per a JSON Schema via response_format; see the Hunyuan invocation guide
  • Image and video: Multimodal understanding · Video understanding; the Hunyuan language models in the table above do not accept image / video input, so use these vision models for multimodal scenarios

Hunyuan vision models​

  • hy-vision-2.0-instruct: A fast-thinking image-to-text model for general image-to-text scenarios. Compared with the previous generation, it delivers notable improvements in basic perception, content recognition, knowledge, OCR, STEM, reasoning, and chart understanding.
  • hunyuan-t1-vision-20250916: A deep-thinking image-to-text model for general image-and-text Q&A, visual grounding, OCR, charts, question solving from photos, and image-based creation. English and minor-language capabilities are also improved.
  • hunyuan-turbos-vision-video-20250728: A video understanding model that supports basic video understanding capabilities such as video description and video content Q&A.

DeepSeek​

Official documentation: DeepSeek API docs

ModelContextImageVideoFileDeep ThinkingTool CallingStructured OutputCaching
deepseek-v4.1-flash1M✅——✅✅✅✅
deepseek/deepseek-flash1M✅——✅✅✅✅
deepseek-v4-pro1M❌——✅✅✅✅
deepseek-v4-pro-08131M❌——✅✅✅✅

Of these, deepseek/deepseek-flash corresponds to flash-v4.1; deepseek-v4-pro is the preview version and deepseek-v4-pro-0813 is the general-availability version.

Zhipu AI​

Official documentation: Zhipu AI documentation

ModelContextImageVideoFileDeep ThinkingTool CallingStructured OutputCaching
glm-5.3-flashx1M✅✅✅✅✅✅✅
glm-5.3-flash1M✅✅✅✅✅✅✅
glm-5.31M———✅✅✅✅
glm-5.21M———✅✅✅✅
  • Vision and files: GLM-5.3-Flash / FlashX; only these two models currently accept document file input, and only via URL
  • Deep thinking: thinking.type only accepts enabled; reasoning_effort accepts low / high / max
  • Structured output: Structured Output

Moonshot AI​

Official documentation: Kimi Chat Completions API

ModelContextImageVideoFileDeep ThinkingTool CallingStructured OutputCaching
kimi-k31M✅✅—✅✅✅✅
kimi-k2.8-preview1M✅✅—✅✅✅✅
kimi-k2.7-code-highspeed256K✅✅—✅✅✅✅
kimi-k2.7-code256K✅✅—✅✅✅✅
kimi-k2.6256K✅✅—✅✅✅✅
  • Vision: Kimi vision models
  • Deep thinking: kimi-k3 always reasons; top-level reasoning_effort accepts low / high / max (default max). K2.x uses thinking.type (enabled / disabled); kimi-k2.7-code cannot disable thinking
  • Structured output: response_format; see Chat Completions API

MiniMax​

Official documentation: Model invocation

ModelContextImageVideoFileDeep ThinkingTool CallingStructured OutputCaching
minimax-m31M✅✅—✅✅—✅
minimax-m2.7204.8K———✅✅—✅
  • Vision: OpenAI-compatible API
  • Deep thinking: thinking.type accepts disabled / adaptive; MiniMax-M3 can disable thinking, M2.x cannot

Xiaomi MiMo​

Official documentation: MiMo open platform docs

ModelContextImageVideoFileDeep ThinkingTool CallingStructured OutputCaching
mimo-v2.6-pro1M✅✅—✅✅✅✅
mimo-v2.6-flash1M✅✅—✅✅✅✅
  • Model details: MiMo-V2.6-Pro, covering multimodal understanding, deep thinking, and structured output parameters

StepFun​

Official documentation: StepFun open platform docs

ModelContextImageVideoFileDeep ThinkingTool CallingStructured OutputCaching
step-5-preview1M✅✅—✅✅✅✅
  • Model details: Step 5 Preview, covering multimodal understanding and JSON Mode / JSON Schema
  • Deep thinking: reasoning_effort accepts low / medium / high

Invocation Methods​

CloudBase offers several ways to call the models, covering Mini Program, Web, server-side, and third-party SDKs. Pick the one that fits your application; detailed usage of each capability lives on the corresponding page under "Feature Guide".

MethodDescription
Mini ProgramWeChat Mini Program, via wx.cloud.extend.AI; no SDK installation required
wx-server-sdkWeChat Cloud Functions, via the built-in wx-server-sdk
Web SDKBrowser / web apps, via @cloudbase/js-sdk
Node SDKNode.js server side (Cloud Functions, Cloud Run), via @cloudbase/node-sdk
cURLDirect HTTP API access for backend services, scripts, and any programming language
OpenAI SDKReuse your existing OpenAI SDK for easy migration and multi-model switching
Anthropic SDKReuse your existing Anthropic SDK — just replace authToken and baseURL
note

The Mini Program, wx-server-sdk, Web SDK, and Node SDK methods authenticate through your CloudBase login state and environment, while cURL, the OpenAI SDK, and the Anthropic SDK require an API Key first.

FAQ​

The same capability takes different parameter names across models. How do I keep this consistent?

CloudBase uses the OpenAI Chat Completions-compatible protocol throughout (Responses API and Anthropic Messages API are also supported), so message structure, tools, image_url, and similar fields are identical on CloudBase. Vendor-specific parameters — such as Hunyuan's thinking levels or Zhipu's thinking.type — follow the relevant official documentation and can be passed through as needed.

Capabilities match the official service — does that mean this page replaces the vendor docs?

No. This page answers "model selection — which models exist and which one can do what"; the vendor documentation answers "implementation — which parameters to set and what the limits are". Use them together: pick the model here, then confirm the parameter details in the vendor documentation.

If a capability is marked —, does that mean it is unsupported?

Not necessarily. — only means the vendor does not state that capability for that specific model, which is common for models the vendor positions as text-only. To determine whether a capability is available, check the model's official documentation, or simply call it from the console to verify.