Release notes

What ships in Apertis: new models, features and fixes. Follow along with the RSS feed.

Category
  1. v2.2.107Feature

    Models Added

    Add GPT-6.1 Sol

    GPT-6.1 Sol

    GPT-6.1 Sol is an upgraded high-end model in OpenAI's GPT-6 series, positioned below the flagship GPT-6 Astra. It is optimized for agentic coding, computer use, document-heavy professional work, and multi-step business automation, delivering near-Astra-level capability at significantly lower cost.

    Compared with GPT-6 Sol, it offers improved factual reliability and stronger adherence to explicit constraints and user intent, making it well suited for complex, long-running agentic workflows where accurate and dependable execution is critical.

  2. v2.2.106Feature

    Models Added

    Add Claude Sonnet 5.5

    Claude Sonnet 5.5

    Claude Sonnet 5.5 is Anthropic's Sonnet-class model for well-scoped everyday work, serving as a direct upgrade to Sonnet 5. It excels at feature development, bug fixing, and creating polished documents, presentations, and spreadsheets, while offering clearer writing and communication than its predecessor.

  3. v2.2.105Feature

    Models Added

    Add GPT-6 Sol & GPT-6 Luna

    GPT-6 Sol

    GPT-6 Sol is OpenAI's cost-efficient high-end model in the GPT-6 series, positioned between the flagship GPT-6 Astra and the fast GPT-6 Luna tier. It is designed for professional knowledge work, agentic coding, business workflow automation, and computer-use tasks, with particular strength in long-horizon software engineering across real-world codebases.

    GPT-6 Sol approaches Astra-level factual reliability at a significantly lower cost, while sharing its clear and concise communication style. This balance of capability, reliability, and efficiency makes it well suited for production agents, complex engineering workflows, and scalable professional applications

    GPT-6 Luna

    GPT-6 Luna is the fast, cost-efficient model in OpenAI's GPT-6 series, optimized for high-volume and latency-sensitive workloads such as chat, classification, and lightweight agentic tasks. It combines low-cost, responsive inference with the GPT-6 family’s improvements in factual reliability and clear, concise communication.

    At higher reasoning effort, GPT-6 Luna can also handle complex software engineering and computer-use workflows that previously required a Sol-tier model, making it a versatile choice for scalable production applications that need to balance speed, cost, and capability.

  4. v2.2.104Feature

    Models Added

    Add Claude Opus 5.5

    Claude Opus 5.5

    Claude Opus 5.5 is Anthropic's flagship model for advanced reasoning, coding, and long-horizon agentic workflows, succeeding Opus 5. It excels at multi-step changes across large codebases, code review and bug detection, financial and scientific analysis, and understanding dense charts, diagrams, and screenshots, with stronger grounding when reporting figures and citing sources. Compared with Opus 5, it completes comparable tasks with fewer steps and lower token usage while providing clearer, more concise progress reporting.

    With adaptive thinking and configurable effort levels, Opus 5.5 can balance reasoning depth, latency, and cost, making it well suited for both demanding autonomous workflows and latency-sensitive professional tasks.

  5. v2.2.103Feature

    Models Added

    Add Grok 4.7

    Grok 4.7

    Grok 4.7 is SpaceXAI's flagship model for coding, agentic workflows, and professional knowledge work, succeeding Grok 4.6. It is particularly strong at long-running software engineering, self-verification, and long-context execution, while improving capabilities in document drafting, presentations, and other professional tasks.

    Trained with extended reinforcement learning focused on multi-hour problems, Grok 4.7 is optimized for sustained, complex task execution and natively supports the Grok Bot harness for conversational workflows. It also introduces an enhanced safeguard stack designed to combine strong jailbreak resistance with low refusal rates for legitimate technical work. Reported benchmark results use xhigh reasoning effort.

  6. v2.2.102Feature

    Price Updated

    Add Credits purchases now carry a volume discount of 5%, 10% or 15%.

    Buying Apertis credits in a single larger purchase now costs less.

    | Add Credits amount | Discount | |:--|--:| | US$15 – 99 | List price | | US$100 – 249 | 5% off | | US$250 – 999 | 10% off | | US$1,000 – 25,000 | 15% off |

    The discount comes off what you pay at checkout. The balance added to your account is always the full amount you selected. Each tier applies to a single purchase; separate purchases are not combined.

  7. v2.2.101Feature

    Models Added

    Add DeepSeek V4.1 Flash

    DeepSeek V4.1 Flash

    DeepSeek V4.1 Flash is a cost-efficient sparse Mixture-of-Experts (MoE) model in DeepSeek's V4.1 family, optimized for coding, reasoning, and agentic workflows. Despite its efficiency-focused positioning, DeepSeek reports that it surpasses the previous V4 Pro in performance, inference speed, and overall task completion time.

    The model is particularly strong at long-horizon, multi-step execution, making it well suited for coding agents, complex problem solving, and autonomous workflows that must reliably carry tasks through to completion.

  8. v2.2.100Feature

    Models Added

    Add GPT-6 Astra

    GPT-6 Astra

    GPT-6 Astra is OpenAI's flagship model for demanding end-to-end professional work, designed for advanced analysis, software engineering, deep research, scientific tasks, and document creation.

    It is particularly strong in long-horizon agentic workflows, including tasks that require sustained reasoning, tool orchestration, and computer and browser use, making it well suited for complex autonomous workflows and production-grade knowledge work.

  9. v2.2.99Feature

    Models Added

    Add Muse Spark 1.3

    Muse Spark 1.3

    Muse Spark 1.3 is Meta's multimodal reasoning model designed for long-running agentic, multi-agent, and coding workflows. It maintains context and information across extended tasks, enabling reliable execution in complex, multi-step environments.

    The model is optimized to resolve conflicting information, seek clarification or confirmation when necessary, and execute concisely, making it well suited for autonomous agents, collaborative multi-agent systems, and long-horizon software engineering workflows.

  10. v2.2.98Feature

    Models Added

    Add Gemini 3.8 Flash

    Gemini 3.8 Flash

    Gemini 3.8 Flash is Google's most intelligent Flash-class model, delivering significant improvements over Gemini 3.7 Flash across software engineering, agentic workflows, and complex multi-step reasoning.

    Designed to combine strong capability with Flash-tier efficiency, it is well suited for coding assistants, autonomous agents, and high-throughput production workflows that require responsive performance without sacrificing reasoning quality.

  11. v2.2.97Feature

    Models Added

    Add Claude Fable 5.1

    Claude Fable 5.1

    Claude Fable 5.1 is an upgraded version of Fable 5, delivering broad improvements with particularly strong gains in agentic coding, long-running workflows, and professional knowledge work. It excels at large code refactors, front-end and visual code generation, financial analysis, and complex analytical tasks.

    Compared with Fable 5, it also produces more concise plans and summaries while maintaining strong performance across extended tasks, making it a natural upgrade for existing Fable workflows and a strong option alongside Opus 5 for reasoning-intensive applications.

  12. v2.2.96Feature

    Models Added

    Add Hy4 preview

    Hy4 preview

    Tencent Hy4 Preview is a Mixture-of-Experts (MoE) model from Tencent, featuring 770B total parameters with 49B activated per token. It is designed for coding agents, complex tool-driven workflows, and professional productivity tasks that require strong planning and reliable execution.

    Optimized for context continuity and sustained multi-step work, Hy4 Preview is well suited for long-horizon coding, agentic automation, tool orchestration, and complex real-world workflows.

  13. v2.2.95Feature

    Models Added

    Add GLM 5.3 Flash

    GLM 5.3 Flash

    GLM-5.3-Flash is Z.ai's efficient native multimodal model, designed for coding and long-horizon agentic workflows. It combines strong multimodal capabilities with an architecture optimized for responsive, cost-efficient task execution.

    Built on a hybrid sparse and linear attention architecture, GLM-5.3-Flash maintains accurate long-context behavior while reducing computational overhead, making it well suited for coding agents, extended multi-step tasks, and scalable production workloads.

  14. v2.2.94Feature

    Models Added

    Add GLM 5.3

    GLM 5.3

    GLM-5.3 is Z.ai's large-scale reasoning model designed for complex software engineering and long-horizon agentic workflows. It supports text input and output with a 1M-token context window, enabling sustained reasoning across large codebases and extended multi-step tasks.

    Building on GLM-5.2, it delivers stronger coding performance while improving the balance between capability and token efficiency, making it well suited for autonomous coding agents, large-scale engineering workflows, and complex task execution.

  15. v2.2.93Feature

    Models Added

    Add Qwen3.8 27B

    Qwen3.8 27B

    Qwen3.8 27B is an open-weight dense vision-language model from Qwen, designed for coding, professional knowledge work, research, and multimodal interaction. It combines strong text and visual understanding with capabilities optimized for sustained, real-world agentic tasks.

    The model supports flexible thinking modes that can be enabled for deeper reasoning or disabled for faster execution, making it well suited for long-running agents, multimodal workflows, coding assistants, and cost-conscious self-hosted deployments.

  16. v2.2.92Feature

    Models Added

    Add Gemini 3.7 Flash

    Gemini 3.7 Flash

    Gemini 3.7 Flash is Google's fast multimodal model designed for agentic workflows, coding, and complex multi-step reasoning. It combines responsive inference with reliable problem-solving capabilities, making it well suited for interactive and production-scale applications.

    Optimized for speed and dependable multi-step execution, Gemini 3.7 Flash is a strong choice for coding assistants, autonomous agents, and high-throughput workflows that require both low latency and capable reasoning.

  17. v2.2.91Feature

    Models Added

    Add Grok 4.6, DeepSeek V4 Pro 0813 & Qwen3.8 2.4T A95B

    Grok 4.6

    Grok 4.6 is SpaceXAI's smartest frontier model, delivering top-tier performance across coding, knowledge work, and STEM reasoning. It is designed for demanding technical and professional workloads that require strong problem solving, accurate instruction following, and reliable execution.

    Optimized for software engineering, scientific analysis, and complex knowledge tasks, Grok 4.6 is well suited for advanced coding, research, and agentic workflows where high capability and reasoning quality are critical.

    DeepSeek V4 Pro 0813

    DeepSeek V4 Pro 0813 is DeepSeek's large-scale Mixture-of-Experts (MoE) model and the general availability (GA) release of DeepSeek V4 Pro. It is designed for high-capability workloads requiring advanced reasoning, coding, and agentic task execution.

    As the production-ready V4 Pro release, it is well suited for complex software engineering, long-horizon agent workflows, and demanding reasoning tasks where reliability and model capability are critical.

    Qwen3.8 2.4T A95B

    Qwen3.8 2.4T A95B is Qwen's open-weight sparse Mixture-of-Experts (MoE) model and the open-weight counterpart to Qwen3.8 Max. It features 2.4T total parameters with 95B activated per token, combining frontier-scale capacity with efficient sparse inference.

    Designed for coding, research, complex reasoning, and agentic workflows, the model is well suited for demanding long-horizon tasks and advanced autonomous systems while providing the flexibility and customization benefits of open weights.

  18. v2.2.90Feature

    Models Added

    Add Nemotron 3.5 Lightning

    Nemotron 3.5 Lightning

    NVIDIA Nemotron 3.5 Lightning is an open Mixture-of-Experts (MoE) model with 30B total parameters and 3B active per token, optimized for high-throughput agentic workloads and efficient inference.

    Its lightweight active compute and open design make it well suited for specialized agents, domain-specific customization, and scalable production deployments where speed, cost efficiency, and adaptability are key.

  19. v2.2.89Feature

    Models Added

    Add Muse Glimmer 30B

    Muse Glimmer 30B

    Muse Glimmer 30B is a dense, open-weight multimodal model from Meta Superintelligence Labs, distilled from Muse Spark and optimized for autonomous agents on consumer hardware. It combines strong multi-step reasoning, reliable tool use, failure recovery, image understanding, and multilingual support across 100+ languages.

    Designed for long-horizon agentic and coding workflows, Muse Glimmer 30B offers a practical balance of capability and deployment efficiency, making it well suited for local coding assistants, multimodal agents, and production workflows that require sustained autonomous execution.

  20. v2.2.88Feature

    Models Added

    Add Muse Spark 1.2

    Muse Spark 1.2

    Muse Spark 1.2 is Meta's multimodal reasoning model designed for complex agentic and software engineering workflows. It supports text, image, video, audio, and PDF inputs with text output, and features a 1M-token context window for sustained reasoning across large, multi-stage tasks.

    Built for flexible multi-agent execution, Muse Spark 1.2 can serve as either a coordinating main agent or a parallel task-focused subagent. With configurable reasoning effort, structured outputs, parallel function calling, and broad coding-harness compatibility, it is well suited for multi-file refactoring, extended debugging, whole-repository generation, and long-horizon development workflows.

    Enjoy.

  21. v2.2.87Feature

    Models Added

    Add Qwen3.8 Max

    Qwen3.8 Max

    Qwen3.8 Max is the flagship model in Alibaba’s Qwen3.8 series and the general-availability successor to Qwen3.8 Max Preview. It is a multimodal reasoning model designed for complex tasks across reasoning, visual understanding, coding, and agentic workflows.

    As the production-ready top tier of the Qwen3.8 family, it is well suited for advanced problem solving, multimodal analysis, software engineering, and long-running tool-driven applications.

    Enjoy.

  22. v2.2.86Feature

    Models Added

    Price Cut for OpenAI GPT-5.6 Terra & Luna

    Price Cut on GPT-5.6 Terra & Luna

    Today, OpenAI made GPT-5.6 more affordable and faster. GPT-5.6 Luna now costs 80% less, GPT-5.6 Terra costs 20% less.

    Therefore, we also lower the prices for GPT-5.6 Terra & Luna models.

    Enjoy.

  23. v2.2.85Feature

    Models Added

    Add Claude Opus 5 (Fast)

    Claude Opus 5 (Fast)

    This is the fast version of Opus 5 model.

    Enjoy.

  24. v2.2.84Feature

    Models Added

    Add Qwen3.7 Flash

    Qwen3.7 Flash

    Qwen3.7 Flash is Alibaba's vision-language reasoning model, designed for multimodal agents, visual coding, search, and computer interaction. It combines fast inference with strong visual understanding, including object recognition, spatial reasoning, and real-world scene perception.

    Optimized for interactive and agentic workflows, Qwen3.7 Flash is well suited for GUI understanding, visual question answering, multimodal search, and computer-use applications that require responsive reasoning across text and images.

    Enjoy.

  25. v2.2.83Feature

    Models Added

    Add Claude Opus 5

    Claude Opus 5

    Claude Opus 5 is Anthropic's flagship model for advanced reasoning, coding, and long-horizon agentic workflows. It excels at end-to-end software engineering, code review, bug detection, visual analysis of charts and documents, complex office deliverables, and parallel subagent coordination.

    The model maintains reliable instruction following and tool use across extended tasks, while remaining effective at lower reasoning-effort settings for workloads that prioritize latency and token efficiency.

    Enjoy.

  26. v2.2.82Feature

    Models Added

    Add Gemini 3.6 Flash & Gemini 3.5 Flash-Lite

    Gemini 3.6 Flash

    Gemini 3.6 Flash is Google's high-efficiency model for coding, agentic workflows, and web and application development. It is optimized to produce polished, production-ready outputs with fewer unnecessary revisions, less hedging, and more direct task execution.

    By reducing both token usage and the number of model calls required to complete complex tasks, Gemini 3.6 Flash is well suited for high-throughput development, scalable agent systems, and cost-sensitive production workflows.

    Gemini 3.5 Flash-Lite

    Gemini 3.5 Flash-Lite is Google's high-efficiency model with enhanced agentic capabilities, optimized for fast, cost-effective inference. It is designed to handle focused tasks with low latency while maintaining strong reasoning and execution quality.

    Well suited for subagents in complex multi-agent systems, Gemini 3.5 Flash-Lite excels at executing specialized tasks within larger workflows, making it ideal for scalable agent orchestration and high-throughput production environments.

    Enjoy them.

  27. v2.2.81Feature

    Models Added

    Add Kimi K3

    Kimi K3

    Kimi K3 is Moonshot AI's 2.8T-parameter open-weight multimodal reasoning model, designed for complex coding, knowledge work, and long-horizon agentic workflows. It excels at repository-scale development, tool use, debugging, and iterative problem solving across text, images, logs, tests, and runtime feedback.

    Built with KDA and Attention Residuals for improved computational efficiency, Kimi K3 delivers strong performance on advanced engineering and multimodal reasoning tasks, making it well suited for autonomous coding agents and large-scale production workflows.

    Enjoy it.

  28. v2.2.80Feature

    Models Added

    Add GPT-5.6 Series

    GPT-5.6 Sol

    GPT-5.6 Sol is the flagship model in OpenAI's GPT-5.6 series, designed for complex reasoning, coding, and agentic workflows. It delivers strong performance on multi-step software engineering tasks, command-line workflows, and long-horizon problem solving, making it well suited for advanced development and autonomous execution.

    Optimized for high-reliability reasoning and end-to-end task completion, GPT-5.6 Sol excels in coding, tool-driven automation, and large-scale engineering workflows that require sustained context and precise execution.

    GPT-5.6 Terra

    GPT-5.6 Terra is the balanced model in OpenAI's GPT-5.6 series, positioned between the flagship Sol tier and the cost-efficient Luna tier. It is designed for everyday coding, reasoning, and agentic workflows, delivering strong performance while balancing capability and cost.

    Offering near-flagship quality at approximately half the cost of Sol, GPT-5.6 Terra is well suited for production applications that require reliable reasoning, software development, and scalable agent execution.

    GPT-5.6 Luna

    GPT-5.6 Luna is the fast, cost-efficient model in OpenAI's GPT-5.6 series, optimized for high-volume, latency-sensitive workloads. It delivers capable reasoning at an affordable price point, making it ideal for chat applications, classification, and lightweight agentic workflows.

    Designed for scalable production deployments, GPT-5.6 Luna balances speed, cost, and reliability, providing efficient performance for real-time applications and large-scale automation tasks.

    Have fun.

  29. v2.2.79Feature

    Models Added

    Add Grok 4.5

    Grok 4.5

    Grok 4.5 is SpaceXAI’s flagship frontier model, delivering top-tier performance across coding, knowledge work, and STEM reasoning. Designed for demanding professional and technical workloads, it combines strong reasoning, accurate instruction following, and robust problem-solving capabilities.

    Optimized for software engineering, scientific analysis, and complex knowledge tasks, Grok 4.5 is well suited for advanced coding, research, and agent-driven workflows that require high accuracy and reliable long-horizon reasoning.

    Enjoy it.

  30. v2.2.78Feature

    Models Added

    Add Hy3

    Hy3 (Free)

    Hy3 is Tencent's 295B-parameter Mixture-of-Experts (MoE) model, activating 21B parameters per token across 192 experts, and designed for reasoning, agentic workflows, and production-scale applications. It supports a 256K-token context window and configurable reasoning modes, including no-think, low, and high reasoning effort to balance speed and problem-solving depth.

    Optimized for long-horizon tasks, coding, and tool-driven execution, Hy3 delivers strong performance in multi-turn reasoning, constraint tracking, and stable tool calling. With an emphasis on grounded responses and reduced hallucinations, it is well suited for software development, document processing, financial analysis, game development, and enterprise agent workflows.

    Enjoy it.

  31. v2.2.77Feature

    Models Added

    Claude Fable 5 is back

    • Claude Fable 5
  32. v2.2.76Feature

    Models Added

    Add Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image)

    Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image)

    Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image) is Google's fastest and most cost-efficient multimodal image generation model, designed for high-throughput visual workflows and real-time applications. It supports text-to-image generation, image editing, and multi-image composition through a unified API, while also producing text outputs alongside images.

    Delivering image generation in approximately 4 seconds, it combines fast inference with strong character consistency, precise editing, and real-world knowledge. The model generates 1K-resolution images across 14 aspect ratios and embeds an invisible SynthID watermark in all outputs. Optimized for the best balance of quality, speed, and cost, Nano Banana 2 Lite is ideal for prototyping, developer pipelines, and large-scale visual content generation.

    Enjoy it.

  33. v2.2.75Feature

    Models Added

    Add Claude Sonnet 5

    Claude Sonnet 5

    Sonnet 5 is Anthropic's most capable Sonnet-class model, delivering frontier-level performance across coding, agentic workflows, and professional knowledge tasks. It supports text, image, and file inputs, features a 1M-token context window, and offers adaptive thinking with configurable reasoning levels (low, medium, high, max, and x-high) to balance speed, cost, and reasoning depth.

    Optimized for complex coding, long-horizon agent execution, and professional workflows, Sonnet 5 combines strong reasoning, robust instruction following, and enhanced safety features, including an updated tokenizer and real-time cyber safeguards for high-risk dual-use scenarios.

    Enjoy it.

  34. v2.2.74Feature

    Models Added

    Add Fugu Ultra

    Fugu Ultra

    Fugu Ultra is the high-performance model in Sakana AI's Fugu family, built as a learned multi-agent orchestration system rather than a single monolithic model. It intelligently routes tasks across a pool of underlying models and can recursively invoke itself to solve complex problems more effectively.

    Optimized for multi-step reasoning, coding, and agentic workflows, Fugu Ultra supports configurable reasoning effort, native tool calling, and built-in web search. Its orchestration-based design makes it well suited for advanced autonomous agents and complex task execution requiring adaptive model coordination.

    Enjoy it.

  35. v2.2.73Feature

    Models Added

    Add GLM 5.2

    GLM 5.2

    GLM-5.2 is Z.AI's flagship model for long-horizon task execution, designed to handle complex, project-scale workflows with high reliability. Featuring a 1M-token context window, it can maintain and reason over extensive engineering context, enabling consistent execution across large, multi-stage tasks.

    Optimized for end-to-end software development, GLM-5.2 follows engineering standards reliably and can manage the full workflow from requirements analysis and implementation to testing and multi-platform deployment, making it well suited for advanced coding agents and large-scale autonomous engineering projects.

    Enjoy it.

  36. v2.2.72Feature

    Models Added

    Add Kimi K2.7 Code

    Kimi K2.7 Code

    Kimi K2.7 Code is a coding-focused model in Moonshot AI's Kimi K2 family, designed for long-horizon software engineering and agentic development workflows. Built on a native multimodal Mixture-of-Experts (MoE) architecture, it supports text, image, and video inputs and operates exclusively in thinking mode, preserving reasoning across multi-turn interactions.

    With approximately 1T total parameters and 32B activated per token, plus a 256K-token context window, K2.7 Code excels at end-to-end programming tasks, agentic task decomposition, repository-scale reasoning, and extended coding conversations, making it well suited for advanced coding agents and long-context development workflows.

    Enjoy it.

  37. v2.2.71Feature

    Models Added

    Add Claude Fable 5

    Claude Fable 5

    Claude Fable 5 is Anthropic's Mythos-class model, designed for autonomous knowledge work, coding, and long-running agentic workflows. It supports text, image, and file inputs with text output, includes reasoning capabilities, and features a 1M-token context window for handling complex, high-context tasks.

    Optimized for asynchronous and long-horizon execution, Claude Fable 5 excels at end-to-end tasks that would typically require hours, days, or weeks of human effort. It combines strong reasoning, autonomous verification and self-correction loops, and robust safeguards, making it well suited for complex research, software engineering, and large-scale knowledge work.

    Enjoy it.

  38. v2.2.70Feature

    Models Added

    Add Nemotron 3 Ultra & Nemotron 3.5 Content Safety

    Nemotron 3 Ultra

    NVIDIA Nemotron 3 Ultra is an open frontier reasoning and orchestration model featuring a 550B-parameter Mixture-of-Experts (MoE) architecture with 55B active parameters per token. Built on a hybrid Transformer–Mamba design, it supports text input and output with a 1M-token context window, enabling large-scale reasoning and long-horizon task execution.

    Optimized for agent orchestration, coding agents, deep research, and complex enterprise workflows, the model excels at multi-step reasoning, planning, and sustained execution. With high-throughput inference designed for large-scale agent pipelines, Nemotron 3 Ultra serves as a powerful foundation for advanced agentic AI systems.

    Nemotron 3.5 Content Safety

    NVIDIA Nemotron 3.5 Content Safety is a compact 4B-parameter multimodal guardrail model from NVIDIA, designed for content moderation, safety classification, and AI policy enforcement. Supporting text and image inputs with text output, it evaluates both user prompts and model responses, providing safe/unsafe classifications, safety category labels, and optional reasoning traces.

    Fine-tuned from Gemma-3-4B and supporting 12 languages with a 128K-token context window, the model is well suited for prompt moderation, response filtering, content classification, and enterprise safety pipelines. As part of the NVIDIA Nemotron family, it offers a configurable reasoning mode and integrates easily into agentic AI systems requiring robust guardrails and compliance controls.

    Enjoy them.

  39. v2.2.69Feature

    Models Added

    Add Qwen3.7 Plus

    Qwen3.7 Plus

    Qwen3.7-Plus is a cost-effective multimodal model in Alibaba's Qwen3.7 series, supporting text and image inputs with text output. It combines the series' strong language capabilities with significantly enhanced vision-language understanding, while retaining full-stack agent-level intelligence for coding, tool use, and productivity workflows.

    Its standout capability is multimodal interactive agency—the ability to perceive real-world scenes, understand screens and graphical interfaces, generate code from visual references, and perform end-to-end navigation within applications. This makes Qwen3.7-Plus well suited for GUI automation, visual coding, productivity agents, and multimodal task execution.

    Enjoy it.

  40. v2.2.68Feature

    Models Added

    Add MiniMax M3

    MiniMax M3

    MiniMax-M3 is a multimodal foundation model from MiniMax, supporting text, image, and video inputs with text output and a 1M-token context window. It is designed for long-horizon agentic workflows, coding, and tool-driven task execution, enabling sustained reasoning across complex tasks.

    Built on MiniMax Sparse Attention (MSA), the model dramatically improves long-context efficiency by replacing full attention with KV-block selection, reducing compute costs at 1M-token contexts while maintaining strong performance. Trained as a native multimodal model and optimized for multi-turn, production-style collaboration, MiniMax-M3 excels at extended, multi-step workflows rather than single-turn interactions.

    Enjoy it.

  41. v2.2.67Feature

    Models Added

    Add Claude Opus 4.8 Series

    Claude Opus 4.8

    Claude Opus 4.8 is Anthropic's most capable generally available model in the Opus family, designed for highly autonomous agents, long-horizon workflows, and advanced knowledge work. It supports text, image, and file inputs with text output, includes reasoning capabilities, and features a 1M-token context window for maintaining coherence across extended tasks and sessions.

    The model excels at multi-step reasoning, complex coding, and end-to-end project orchestration, including large codebases, multi-stage debugging, and long-running asynchronous agent pipelines. Beyond software engineering, it is highly effective for document drafting, presentation creation, data analysis, and memory-driven workflows, delivering consistent quality across very long outputs and complex projects.

    Enjoy it.

  42. v2.2.66Feature

    Models Added

    Add Qwen3.7 Max

    Qwen3.7 Max

    Qwen3.7-Max is the flagship model in Alibaba's Qwen3.7 series, designed for agent-centric workloads with strong performance in coding, productivity, and long-horizon autonomous execution. It supports text input and output and delivers notable improvements in coding and agentic capabilities over previous Qwen generations.

    Optimized for real-world workflows, the model also supports explicit prompt caching for efficient reuse of repeated context, making it well suited for scalable development, office automation, and advanced agent systems.

    Enjoy it.

  43. v2.2.65Feature

    Models Added

    Add Grok Build 0.1

    Grok Build 0.1

    Grok Build 0.1 is xAI's fast coding model designed specifically for agentic software engineering workflows. It supports text and image inputs with text output, and is optimized for interactive coding agents, tool use, and multi-step development tasks.

    Powering the Grok Build CLI, the model features a 256K token context window with effectively no text output limit, making it well suited for long-horizon coding, automation, and continuous development workflows. Currently available in early access.

    Enjoy it.

  44. v2.2.64Feature

    Feature Added

    Apertis Coworker — delegate grunt work to cheaper models

    The Apertis MCP server now ships a delegate coworker tool. Connect it to Claude Code (or any MCP client) and Claude can hand off routine, high-volume subtasks — bulk edits, boilerplate, repetitive lookups — to a cheaper model through your Apertis API key, while staying in control as the "manager."

    What's new
    • delegate tool — Claude calls a single MCP tool to run a subtask on a lower-cost model, then reviews the result. You keep premium models for reasoning and spend cheap tokens on the busywork.
    • Works with any MCP client — Claude Code, Claude Desktop, and other MCP-compatible agents.
    • One install — available on npm as @apertis/mcp-server.
    Get started
    npx @apertis/mcp-server
    

    Add it to your MCP client config with your Apertis API key, and Claude can start delegating immediately.

    See the official documentation here and source code.

  45. v2.2.63Feature

    Feature Added

    Switch Advisor — see the cost before you switch models

    New in Settings → Usage

    Switch Advisor estimates what your recent usage would cost on a different model — projected from your own pay-as-you-go activity, not a generic price list.

    How it works

    • Pick a From model — your recently used models, highest-spend preselected
    • Search any platform model as the Candidate
    • Get an instant projection: current cost, projected cost, and savings (absolute + %)

    The projection replays your real input/output token volumes from the selected time range on the candidate model's pricing — so the figure reflects how you actually use the model, not a marketing average.

    Example: a workload on claude-opus-4-6 costing $36.00 over 30 days projects to $2.61 on deepseek-v4-pro — a 92.75% reduction.

    Available now in Settings → Usage, right below the per-model usage table. No setup required.

    ▎ Cost estimate only — it does not compare output quality or response speed.

  46. v2.2.62Feature

    Models Added

    Add Gemini 3.5 Flash

    Gemini 3.5 Flash

    Gemini 3.5 Flash is Google's high-efficiency multimodal model, delivering near-Pro level performance in coding and reasoning at Flash-tier speed and cost. It supports text, image, video, audio, and PDF inputs, making it well suited for diverse multimodal workflows.

    Optimized for coding proficiency and parallel agentic execution, the model defaults to medium thinking effort for faster, cost-efficient responses while supporting configurable thinking levels (minimal, low, medium, high) for fine-grained cost–performance control.

    Enjoy it.

  47. v2.2.61Feature

    Models Added

    Add Claude Opus 4.7 (Fast)

    Claude Opus 4.7 (Fast)

    Fast version of Claude Opus 4.7 is live.

    Enjoy it.

  48. v2.2.60Feature

    System Update

    Model Availability Heartbeat

    What changed
    • Added a Recent Availability card to model detail pages.
    • Added heartbeat bars for the recent delivery window.
    • Added hover and keyboard-focus tooltips with date, time, availability percentage, and status.
    • Counted retried or fallback-routed requests as healthy when the request ultimately succeeds.
    How to read it

    The card reports observed delivery, not a synthetic uptime probe.

    If Apertis has recent successful delivery for a model, the relevant heartbeat bucket is green. If an actual delivery failure is observed, the bucket can move to degraded or unavailable based on the observed success rate.

    When there is no recent traffic for a bucket, Apertis treats that silence as no observed anomaly and displays it as green 100%. This keeps the signal aligned with the rule that a model should not look unhealthy just because no one called it during that interval.

    Why this matters

    You can now check model-level health from the same page where you review pricing, context, endpoints, and examples. Teams choosing between models can see recent delivery quality without waiting for a separate status page or paying for active probes.

    The implementation stays cost-aware by using delivery results Apertis already sees during normal routing.

    Enjoy it.

  49. v2.2.59Feature

    Models Added

    Add Mistral Medium 3.5 & Baidu Cobuddy

    Mistral Medium 3.5

    Mistral Medium 3.5 is a 128B dense instruction-following model from Mistral AI, supporting text and image inputs with text output. It is designed for agentic workflows, coding, and complex multi-step reasoning, with strong reliability in multi-tool orchestration and long-horizon tasks.

    The model features a 256K token context window, configurable reasoning effort per request, and a custom vision encoder that handles variable image sizes and aspect ratios. With support for self-hosting on as few as four GPUs and availability under open weights, it is well suited for scalable, production-grade deployments.

    Cobuddy

    CoBuddy is a code generation model from Baidu, optimized for coding tasks and AI agent workflows. It delivers high inference throughput and low end-to-end latency, making it well suited for responsive development and automation environments.

    The model includes native support for tool calling and reasoning, runs with FP8 quantization for efficient deployment, and supports a 131K token context window with up to 65K output tokens, enabling long-context coding and multi-step agentic workflows.

    Enjoy them.

  50. v2.2.58System

    System Update

    Apertis New Look

    Apertis has received a broad UI/UX refresh across public pages, model browsing, API key selection, and admin workflows. This update focuses on a cleaner visual system, clearer product navigation, and more practical operational screens for daily use.

    Highlights
    • Refreshed the public experience with a redesigned 404 page, improved layout chrome, and a more polished route diagnostic view.
    • Introduced the new Verbatim landing experience, including updated positioning, product sections, code preview, and lead capture flow.
    • Added a signed-in Dashboard shortcut in the header for faster access to the main workspace.
    • Redesigned the chat API key selection modal with a cleaner Apertis-style layout, clearer selection states, and improved preview coverage.
    • Improved model detail pages with task-driven endpoint selection, better TTS examples, and cleaner handling of voice and web-search pricing states.
    • Refined the model filter sidebar so empty task categories are hidden, making model discovery easier to scan. overview card.
    • Shared FAQ row styling across pages for a more consistent Apertis UI system.
    • Cleaned up footer navigation and naming, including the updated Verbatim label.
    User Experience

    The new look reduces visual clutter, improves spacing consistency, and makes important actions easier to find.

    Public pages now feel more aligned with the Apertis brand, while authenticated workflows are more direct and data-focused.

    Quality

    This refresh includes focused test coverage for updated components and route behavior, including header navigation, modal behavior, layout chrome, pricing helpers, filter behavior, and cost analytics logic.

    Enjoy them.

  51. v2.2.57Feature

    System Update

    Audio APIs Now Live

    Full Audio API Support

    Apertis now supports the OpenAI-compatible Audio API. Use a single API key to access leading TTS (text-to-speech) and STT (speech-to-text) models across providers.

    Supported Models

    Text-to-Speech (TTS)

    • gemini-3.1-flash-tts-preview — Google's latest Flash TTS preview
    • gpt-4o-mini-tts — OpenAI's lightweight real-time speech synthesis

    Speech-to-Text (STT)

    • gpt-4o-transcribe — Flagship high-accuracy transcription
    • gpt-4o-mini-transcribe — Cost-efficient real-time transcription
    • whisper-large-v3-turbo — Accelerated Whisper v3
    • whisper-large-v3 — Full-precision Whisper
    • whisper-1 — The classic, battle-tested baseline
    Endpoints

    Drop-in compatible with the OpenAI SDK — no code changes required:

    • POST /v1/audio/speech — text → audio
    • POST /v1/audio/transcriptions — audio → text
    • POST /v1/audio/translations — audio → translated text
    Billing
    • PAYG (pay-as-you-go): shares the same quota balance as chat/completions
    • Per-dimension billing: priced separately on input tokens / output tokens / audio seconds, with admin-tunable AudioRatio
    • File limit: 25 MB per multipart upload
    • Subscriptions: audio models are PAYG-only for now (not included in subscription plans)
    Example
    from openai import OpenAI
    
    client = OpenAI(
        api_key="sk-your-apertis-key",
        base_url="https://api.apertis.ai/v1"
    )
    
    # TTS
    speech = client.audio.speech.create(
        model="gpt-4o-mini-tts",
        voice="alloy",
        input="Hello from Apertis."
    )
    speech.stream_to_file("hello.mp3")
    
    # STT
    with open("audio.mp3", "rb") as f:
        transcript = client.audio.transcriptions.create(
            model="whisper-large-v3-turbo",
            file=f
        )
    print(transcript.text)
    
    Model Detail Page Updates
    • Endpoint and code samples auto-switch based on the model's task
    • TTS models now emit ready-to-run OpenAI SDK Python snippets
    • Web Search pricing column hidden for voice models (:web is unsupported)
  52. v2.2.56Feature

    Models Added

    Add Grok 4.3

    Grok 4.3

    Grok 4.3 is a reasoning-focused model from xAI designed for agentic workflows, instruction following, and high factual accuracy tasks. It supports text and image inputs with text output, with reasoning always active and not configurable by effort level.

    The model features a 1M-token context window with effectively no output token limit, making it well suited for long-document analysis, deep research, and multi-step agentic workflows. It uses tiered pricing, with higher rates applied to requests exceeding 200K total tokens.

    Enjoy it.

  53. v2.2.55Feature

    Models Added

    Add Nemotron 3 Nano Omni (Free)

    Nemotron 3 Nano Omni (Free)

    NVIDIA Nemotron 3 Nano Omni is an open 30B-A3B multimodal model designed as a perception and context sub-agent for enterprise agent systems. It supports text, image, video, and audio inputs with text output, enabling unified multimodal reasoning within a single inference loop. Built on a hybrid MoE Transformer–Mamba architecture with Conv3D video layers and Efficient Video Sampling (EVS), it delivers significantly improved efficiency for video reasoning—achieving ~2× higher throughput and 2.5× lower compute compared to separate pipelines.

    With up to 300K context length and extended thinking support, it is well suited for scalable, multimodal agent workflows.

    Enjoy it.

  54. v2.2.54Feature

    Models Added

    Add latest Qwen Models

    Qwen3.5 Plus 2026-04-20

    Qwen3.5 Plus (April 2026) is a large-scale multimodal language model from Alibaba, supporting text, image, and video inputs with text output. It features a 1M-token context window, enabling large-scale reasoning and multimodal workflows within a single interaction.

    This updated version of Qwen3.5 Plus introduces tiered pricing beyond 256K tokens, making it suitable for high-context applications while maintaining flexibility for cost optimization in long-input scenarios.

    Qwen3.6 Flash

    Qwen3.6 Flash is a fast and efficient model from Alibaba's Qwen 3.6 series, supporting text, image, and video inputs with a 1M-token context window for high-context multimodal workflows.

    Optimized for performance and cost efficiency, it features tiered pricing beyond 256K tokens and supports prompt caching with both cache creation and read pricing, making it well suited for large-scale, high-throughput applications.

    Qwen3.6 Max Preview

    Qwen3.6-Max-Preview is a proprietary frontier model from Alibaba Cloud built on a sparse Mixture-of-Experts (MoE) architecture with approximately 1 trillion parameters. It is optimized for agentic coding, tool use, and long-context reasoning, supporting a 262K token context window.

    The model includes an integrated thinking mode that preserves reasoning across multi-turn interactions, along with support for structured outputs and function calling.

    Enjoy them.

  55. v2.2.53Feature

    Models Added

    Add GPT-5.5 & GPT-5.5 Pro

    GPT-5.5

    GPT-5.5 is OpenAI's frontier model for complex professional workloads, building on GPT-5.4 with stronger reasoning, higher reliability, and improved token efficiency on challenging tasks. It supports text and image inputs and features a 1M+ token context window (≈922K input, 128K output) for large-scale, high-context workflows.

    Designed for advanced applications, GPT-5.5 excels in reasoning, coding, and multimodal workflows, enabling efficient execution of complex, multi-step tasks within a single system.

    GPT-5.5 Pro

    GPT-5.5 Pro is OpenAI's high-capability model optimized for deep reasoning and accuracy on complex, high-stakes workloads. It supports text and image inputs and features a 1M+ token context window (≈922K input, 128K output) for handling large-scale, long-context tasks.

    Designed for long-horizon problem solving, agentic coding, and precise multi-step execution, GPT-5.5 Pro delivers strong reliability and performance across advanced engineering, research, and complex workflow scenarios.

    Enjoy them.

  56. v2.2.52Feature

    Models Added

    Add DeepSeek V4 Pro & DeepSeek V4 Flash

    DeepSeek V4 Pro

    DeepSeek V4 Pro is a large-scale Mixture-of-Experts (MoE) model with 1.6T total parameters and 49B activated per token, supporting a 1M-token context window for advanced reasoning and long-horizon workflows.

    It delivers strong performance across knowledge, mathematics, and software engineering tasks, making it suitable for complex, real-world applications.

    DeepSeek V4 Flash

    DeepSeek V4 Flash is an efficiency-optimized Mixture-of-Experts (MoE) model with 284B total parameters and 13B activated per token, designed for fast inference and high-throughput workloads.

    It supports a 1M-token context window, enabling large-scale reasoning and long-context processing.

    Enjoy them.

  57. v2.2.51Feature

    Models Added

    Add Qwen3.6-35B-A3B & Qwen3.6-27B

    Qwen3.6-35B-A3B

    Qwen3.6-35B-A3B is an open-weight Mixture-of-Experts (MoE) multimodal model designed for agentic coding and long-horizon workflows. It features ~35–36B total parameters with ~3B activated per token, enabling strong performance with high inference efficiency. The model supports text and image inputs with a ~260K token context window, and is optimized for repository-level reasoning, multi-step development, and tool-driven workflows.

    With strong benchmark performance and improved coherence across extended tasks, Qwen3.6-35B-A3B is well suited for developer tools, coding agents, and real-world engineering applications that require both reasoning depth and efficiency.

    Qwen3.6-27B

    Qwen3.6-27B is an open-weight 27B-parameter dense multimodal model from the Qwen3.6 series, designed to deliver flagship-level coding and agentic performance at a practical deployment scale. It supports both text and image inputs and introduces improvements in agentic coding, repository-level reasoning, and iterative development workflows. Despite its relatively compact size, it achieves state-of-the-art results on coding benchmarks, outperforming much larger models in tasks such as SWE-bench and terminal-based workflows.

    It also provides strong reasoning and multimodal capabilities, along with features like thinking preservation to maintain context across interactions, making it well suited for developer tools, coding agents, and real-world engineering tasks.

    Enjoy it.

  58. v2.2.50Feature

    Models Added

    Add Xiaomi MiMo-V2.5 & MiMo-V2.5-Pro

    MiMo-V2.5

    MiMo-V2.5 is Xiaomi's native omnimodal model, delivering pro-level agentic performance at roughly half the inference cost. It surpasses MiMo-V2-Omni in multimodal perception, particularly in image and video understanding. With a 1M-token context window, it can handle complete documents, extended conversations, and complex task contexts in a single pass.

    Combining strong reasoning, rich perception, and cost efficiency, MiMo-V2.5 is well suited for integration into advanced agent frameworks and real-world multimodal applications.

    MiMo-V2.5-Pro

    MiMo-V2.5-Pro is Xiaomi's flagship model, delivering top-tier performance in agentic capabilities, complex software engineering, and long-horizon tasks. It ranks highly on benchmarks such as ClawEval, GDPVal, and SWE-bench Pro, demonstrating strong real-world reliability. The model can autonomously complete professional tasks that would take human experts days or weeks, executing thousands of tool calls within a single workflow.

    With a 1M-token context window, it is well suited for integration into advanced agent frameworks and large-scale task orchestration systems.

    Enjoy it.

  59. v2.2.49Feature

    Models Added

    Add GPT Image 2

    GPT Image 2

    GPT Image 2 combines OpenAI's GPT-5.4 with advanced image generation capabilities from GPT Image 2, enabling fully integrated multimodal workflows.

    It allows users to seamlessly transition between reasoning, coding, and visual generation within a single interaction, making it well suited for creative, development, and agent-driven applications that require both intelligence and visual output.

    Enjoy it.

  60. v2.2.48Feature

    Models Added

    Add Kimi K2.6

    Kimi K2.6

    Kimi K2.6 is Moonshot AI's next-generation multimodal model, designed for long-horizon coding, UI/UX generation, and multi-agent orchestration. It handles complex end-to-end development tasks across languages such as Python, Rust, and Go, and can transform prompts and visual inputs into production-ready interfaces.

    Powered by a scalable agent swarm architecture, K2.6 can coordinate hundreds of parallel sub-agents for autonomous task decomposition, enabling the generation of documents, websites, and spreadsheets in a single run without human intervention.

    Enjoy it.

  61. v2.2.47Feature

    Feature Added

    Skills & MCP Server — `@apertis/mcp-server` v0.3.0

    A Model Context Protocol server that lets any MCP-compatible AI assistant call Apertis directly. No setup, no wrapper code — just install once.

    Install (Claude Code):

    claude mcp add apertis -- npx -y @apertis/mcp-server
    

    Nine tools shipped:

    | Tool | What it does | |------|--------------| | list_models | List models with optional free/paid + capability filters | | get_model_info | Detailed info for a specific model (pricing, context, provider) | | compare_models | Side-by-side comparison of 2–5 models | | check_quota | Account balance, subscription status, remaining quota | | get_usage_stats | Usage by model and period (today / week / month) | | list_api_keys | List your keys (masked) with status and quota | | create_api_key | Create a new key with an optional quota limit | | suggest_model | Freeform keyword search over the full catalog | | recommend_model | Curated Apertis pick for a task type with live pricing (new in v0.3.0) |

    → Guide: docs.apertis.ai/api/sdks/mcp-server → npm: @apertis/mcp-server

    Agent Skills — one-command install for 45+ AI tools

    Three curated skills that teach your AI assistant how to use Apertis correctly. Install once, works everywhere.

    npx skills add theQuert/apertis-skills
    

    Compatible with Claude Code, Cursor, GitHub Copilot, Codex, Gemini CLI, and 45+ other AI coding tools.

    | Skill | What your agent learns | |-------|------------------------| | apertis-api | Auth, endpoints, :web suffix, MCP reference — the complete API surface | | apertis-model-picker | Opinionated model picks by task type with reasoning | | apertis-migrate | One-line swap from OpenAI SDK to Apertis |

    → Source: github.com/theQuert/apertis-skills

    GET /v1/recommend — dynamic model selection endpoint

    Ask Apertis what to use for a task and get back the curated pick with live pricing. Recommendations update as models are added, retired, or re-priced — your code stays the same.

    curl "https://api.apertis.ai/v1/recommend?task=coding&budget=medium" \
      -H "Authorization: Bearer $APERTIS_API_KEY"
    

    Task types: coding, long-context, fast-chat, reasoning, vision Budget tiers: low, medium (default), high

    Response shape:

    {
      "model": "claude-sonnet-4-6",
      "pricing": { "input_per_1m": 2.40, "output_per_1m": 12.00 },
      "why": "Best coding ability per dollar. 200K context.",
      "alternatives": [
        { "model": "deepseek-v3", "note": "3x cheaper, good for simpler coding" },
        { "model": "claude-opus-4-6", "note": "most capable, higher cost" }
      ]
    }
    

    Use the returned model ID directly in your next /v1/chat/completions call.

    → Reference: docs.apertis.ai/api/utilities/recommend


    Docs

    • New: @apertis/mcp-server SDK guide with recommend_model walkthrough
    • New: GET /v1/recommend endpoint reference with Python example
    • Updated: Cursor integration guide with new screenshots and apertis/ prefix convention
    • Updated: Ideas page now publicly browsable at docs.apertis.ai/help/ideas
    • Updated: Timeout documentation — X-Timeout header, 408 status semantics

    Why this release

    We kept seeing two questions in support:

    1. “Which model should I use?”
    2. “How do I wire Apertis into my agent/IDE?”

    recommend_model + the skills + the MCP server answer both — without asking you to paste the same instructions into every new session. Your agent now picks the right model and knows how to call us, natively.

  62. v2.2.46System

    Subscription Quota Multiplier Update

    Quota multiplier adjustments for Lite, Pro, and Plus plans — effective 2026-04-24 00:00 UTC

    Subscription Quota Multiplier Update

    Effective 2026-04-24 at 00:00 UTC, quota multipliers across Lite, Pro, and Plus plans will be adjusted in response to recent upstream AI provider pricing changes.

    Why
    • Anthropic Claude (upstream routing channel) base rates have trended upward across multiple providers
    • Z.AI GLM-5.1 base rates have trended upward across multiple providers
    • OpenAI GPT-5.4 serving costs have increased on several routes
    • Additionally, claude-opus-4-6 on the Lite plan is being lowered because our review showed the previous multiplier was set higher than current real cost justifies
    Changes
    Lite Plan ($12/month, 600 quota per cycle)
    • glm-5.1: 0.51 → 1.5
    • code:claude-opus-4-6: 5.00 → 10.0
    • claude-opus-4-6: 5.00 → 2.0 (decrease)
    • gemini-3-flash-preview: 0.10 → 0.2
    • gemini-3.1-pro-preview: 0.43 → 0.77
    Pro Plan ($25/month, 900 quota per cycle)
    • glm-5.1: 0.37 → 1.5
    • claude-opus-4-6: 1.00 → 1.5
    • gpt-5.4: 0.62 → 1.0
    • code:claude-opus-4-6: 0.75 → 1.5
    Plus Plan ($60/month, 1,500 quota per cycle)
    • claude-opus-4-6: 0.75 → 1.5
    • glm-5.1: 0.25 → 1.0
    • gpt-5.4: 0.45 → 0.8
    • claude-opus-4-7: 3.00 → 4.0
    What stays the same
    • Monthly subscription fee
    • Billing cycle and quota allowance per plan
    • Model access and plan tiers
    • Pay-As-You-Go (PAYG) fallback behavior
    Your options

    If you do not agree with the changes, you may cancel your subscription at any time before 2026-04-24 00:00 UTC at https://apertis.ai/setting. Affected users will also be notified by email.

    For questions, contact us at [email protected].

  63. v2.2.45Feature

    Models Added

    Add Claude Opus 4.7

    Claude Opus 4.7

    Opus 4.7 is the next generation of Anthropic's Opus family, designed for long-running, asynchronous agent workflows. Building on Opus 4.6, it delivers stronger performance on complex, multi-step tasks and more reliable execution across extended pipelines such as large codebases, multi-stage debugging, and end-to-end project orchestration. Beyond coding, Opus 4.7 enhances knowledge work capabilities, including document drafting, presentation creation, and data analysis. With strong coherence over long outputs and sessions, it is well suited for tasks requiring persistence, judgment, and sustained execution.

    Enjoy it.

  64. v2.2.44Feature

    Feature Added

    Fallback Timeout Setting for Coding Plan Users

    What's new

    When Apertis routes your request to an upstream provider, it waits a set amount of time before switching to the next available channel. Previously this was fixed at 30 seconds — fine for most models, but too short for preview and reasoning models processing large context windows.

    You can now adjust this in Settings → Subscription Keys → Fallback Timeout (range: 5s–300s).

    Who should change this
    • Using gemini-3-flash-preview, claude-opus-4-thinking, or other preview/reasoning models with large prompts? Increase to 120s+
    • Using standard models like gpt-4o, claude-sonnet-4? Default 30s is fine
    How it works
    1. Go to Settings → Subscription Keys
    2. Find Fallback Timeout in the metadata section
    3. Enter your preferred value in milliseconds (e.g., 120000 for 120s)
    4. Click Save

    Changes take effect immediately.

  65. v2.2.43Feature

    Models Added

    Add Claude Opus 4.6 (Fast)

    Claude Opus 4.6 (Fast)

    Opus 4.6 is Anthropic's more faster version of Opus 4.6 model for coding and long-running professional workflows, designed for agents that operate across entire workflows rather than single prompts. It demonstrates strong performance on large codebases, complex refactoring, and multi-step debugging, with improved contextual understanding, deeper problem decomposition, and higher reliability on challenging engineering tasks compared to earlier generations. Beyond software development, Opus 4.6 excels at sustained knowledge work, producing near production-ready documents, technical plans, and analyses in a single pass while maintaining coherence across long outputs and extended sessions. Its strength in persistence, judgment, and structured execution makes it well suited for technical design, migration planning, and end-to-end project execution.

    Enjoy it.

  66. v2.2.42Feature

    Models Added

    Add GLM 5.1

    GLM 5.1

    GLM-5.1 delivers a major advancement in coding capability, with significant improvements in handling long-horizon tasks. It is designed to operate beyond short interactions, enabling continuous, autonomous execution over extended periods. The model can work independently on a single task for 8+ hours, performing planning, execution, and iterative self-improvement to produce complete, engineering-grade results, making it well suited for complex development workflows and autonomous agent systems.

    Enjoy it.

  67. v2.2.41Feature

    System Update

    Add local token estimation for Claude Code CLI

    We've just shipped a fix that implements this endpoint. It performs local token estimation, so it returns a valid {"input_tokens": N} response without any additional latency.

    The endpoint supports:

    • String and structured content block messages
    • System prompts
    • Tool definitions

    The fix is resolved and on Live. Claud Code CLI should start normally against the Apertis API.

    Enjoy it.

  68. v2.2.40Feature

    Models Added

    Add Gemma 4 31B & Gemma 4 26B A4B

    Gemma 4 31B

    Gemma 4 31B Instruct is Google DeepMind's 30.7B dense multimodal model, supporting text and image inputs with text outputs. It features a 256K token context window, configurable thinking/reasoning modes, native function calling, and broad multilingual support across 140+ languages. The model delivers strong performance in coding, reasoning, and document understanding, making it well suited for developer workflows, multilingual applications, and structured knowledge tasks.

    Gemma 4 26B A4B

    Gemma 4 26B A4B IT is an instruction-tuned Mixture-of-Experts (MoE) model from Google DeepMind, featuring 25.2B total parameters with only 3.8B activated per token—delivering near 31B-class quality at a fraction of the compute cost. It supports multimodal inputs including text, images, and video (up to 60s at 1fps). The model includes a 256K token context window, native function calling, configurable thinking/reasoning modes, and structured output support. Released under the Apache 2.0 license, it is well suited for efficient, production-ready multimodal and agentic applications.

    Enjoy it.

  69. v2.2.39Feature

    Models Added

    Add GLM 5V Turbo

    GLM 5V Turbo

    GLM-5V-Turbo is Z.ai's first native multimodal agent foundation model, designed for vision-based coding and agent-driven workflows. It natively supports image, video, and text inputs, enabling integrated multimodal reasoning and execution. The model excels at long-horizon planning, complex coding, and multi-step task execution, and works seamlessly with agents to complete the full loop of “perceive → plan → execute”, making it well suited for advanced multimodal automation and real-world agent systems.

    Enjoy it.

  70. v2.2.38Feature

    Models Added

    Add Grok 4.20

    Grok 4.20

    Grok 4.20 is xAI's newest flagship model, designed for high-performance reasoning with industry-leading speed and strong agentic tool-calling capabilities. It emphasizes strict prompt adherence and low hallucination rates, delivering highly precise and reliable responses. Optimized for agent workflows and real-time applications, Grok 4.20 provides consistent, truthful outputs while maintaining fast inference and robust task execution.

    Grok 4.20 Multi-Agent

    Grok 4.20 Multi-Agent is a specialized variant of xAI's Grok 4.20 designed for collaborative, agent-based workflows. It enables multiple agents to operate in parallel, coordinating tool use and synthesizing information to handle complex, multi-step tasks. Optimized for deep research and large-scale problem solving, the model supports configurable reasoning effort: 4 agents for low/medium settings and up to 16 agents for high/xhigh settings, enabling scalable parallel reasoning and execution.

    Enjoy it.

  71. v2.2.37Feature

    Models Added

    Add Qwen3.6 Plus Preview

    Qwen3.6 Plus Preview (Free)

    Qwen 3.6 Plus Preview is the next-generation evolution of the Qwen Plus series, built on an advanced hybrid architecture that enhances efficiency and scalability. It delivers improved reasoning capabilities and more reliable agentic behavior compared to the 3.5 series, with benchmark performance at or above leading state-of-the-art models.

    Designed as a flagship preview model, it excels in agentic coding, front-end development, and complex problem solving, making it well suited for advanced development workflows and high-performance applications.

    Enjoy it.

  72. v2.2.36Feature

    System Update

    🚀 Ideas Board is Now Public

    What's New
    • Public Read Access — Anyone can now browse ideas, view details, and read comments without signing in
    • Write Actions Require Login — Submitting ideas, voting, and commenting still require authentication
    • Privacy Preserved — Non-public ideas remain hidden from unauthenticated visitors
    Why This Change

    We believe in building in public. The Ideas board is where our community shapes the future of Apertis — sharing feature requests, voting on priorities, and discussing what matters most.

    By making it publicly visible, potential users can see what's being built and why, existing users can share links to ideas they care about, and everyone gets a transparent view into our roadmap and decision-making process.

    Great products are built together. Come see what the community is asking for — and when you're ready, sign in to add your voice.

    Happy Building

  73. v2.2.35Feature

    System Update

    ✨ Kilo CLI & Kilo Code — Official Model Provider

    • Apertis is now available as an official built-in model provider in Kilo CLI and Kilo Code. Users can select Apertis directly from the provider list without any manual endpoint configuration.

    Full notes on apertis.ai

  74. v2.2.34Feature

    System Update

    ✨ Apertis Model ID Prefix Support

    We now support an apertis/ namespace prefix for all model IDs, designed specifically for Cursor IDE users who experience model ID collisions with Cursor's built-in model routing.

    How It Works

    When your model ID overlaps with Cursor's built-in models, simply add the apertis/ prefix to ensure requests route through Apertis:

    | Standard ID | Cursor IDE ID | |-------------|---------------| | gpt-5.4 | apertis/gpt-5.4 | | claude-sonnet-4.5 | apertis/claude-sonnet-4.5 | | code:claude-sonnet-4.5 | apertis/code:claude-sonnet-4.5 | | gpt-5.4:web | apertis/gpt-5.4:web |

    The apertis/ prefix is automatically stripped on our end — all downstream processing (model validation, channel routing, billing, and logging) uses the original model ID. Fully compatible with existing suffixes like :web and code: prefix and all models.

    Model Detail Page

    Each model's detail page now includes a Cursor IDE Model IDs section listing all available apertis/-prefixed identifiers. Click to copy, and hover the infoicon for context.

    Happy Building

  75. v2.2.33Feature

    System Update

    ✨ AWS Partner Network Member

    Our AI infrastructure is powered by Amazon Bedrock, providing enterprise-grade reliability and compliance for all users.

    • AWS Partner Network: Registered as Apertis AI (Stima AI LLC), Partner status Active
    • Powered by Amazon Bedrock: Direct access to Claude models (Sonnet, Opus, Haiku) via AWS infrastructure
    • Footer trust badges: Added AWS Partner Network Member and Powered by Amazon Bedrock badges alongside existing OWASP, NIST CSF, and PCI DSS compliance indicators

    Happy Building.

  76. v2.2.32Feature

    Models Added

    Add Xiaomi Models: MiMo-V2-Pro & MiMo-V2-Omni

    MiMo-V2-Pro

    MiMo-V2-Pro is Xiaomi's flagship foundation model with over 1T parameters and a 1M-token context window, optimized for advanced agentic workflows. It is highly adaptable to general agent frameworks such as OpenClaw, delivering strong performance in complex, real-world task execution. Ranking among the top tier on benchmarks like PinchBench and ClawBench, with performance approaching models like Opus 4.6, MiMo-V2-Pro is designed to act as the core intelligence of agent systems, orchestrating workflows, driving production engineering tasks, and delivering reliable results at scale.

    MiMo-V2-Omni

    MiMo-V2-Omni is a frontier omni-modal model that natively processes image, video, and audio inputs within a unified architecture. It combines strong multimodal perception with advanced agentic capabilities, including visual grounding, multi-step planning, tool use, and code execution. With a 256K context window, MiMo-V2-Omni is well suited for complex real-world tasks that span multiple modalities, enabling integrated reasoning and execution across diverse input types.

    Happy Building.

  77. v2.2.31Feature

    Models Added

    Add MiniMax M2.7

    MiniMax M2.7

    MiniMax-M2.7 is a next-generation large language model designed for autonomous, real-world productivity and continuous improvement. It incorporates advanced multi-agent collaboration, enabling the model to plan, execute, and iteratively refine complex tasks across dynamic environments.

    Built for production-grade workflows, M2.7 supports tasks such as live debugging, root cause analysis, financial modeling, and full document generation across Word, Excel, and PowerPoint. With strong benchmark performance—including 56.2% on SWE-Pro, 57.0% on Terminal Bench 2, and 1495 ELO on GDPval-AA—it sets a new standard for multi-agent systems in real-world digital workflows.

    Happy Building

  78. v2.2.30Feature

    Feature Added

    🚀 Apertis SDK v2.1 — Better compatible with OpenCode, Kilo Code & all AI coding tools

    Full notes on apertis.ai

  79. v2.2.29Feature

    Models Added

    Add GPT-5.4 Mini & GPT-5.4 Nano

    GPT-5.4 Mini

    GPT-5.4 mini brings the core capabilities of GPT-5.4 into a faster, more efficient model optimized for high-throughput workloads. It supports text and image inputs and delivers strong performance across reasoning, coding, and tool use, while reducing latency and cost for large-scale deployments.

    Designed for production environments, GPT-5.4 mini balances capability and efficiency, making it well suited for chat applications, coding assistants, and scalable agent workflows. It provides reliable instruction following, solid multi-step reasoning, and consistent performance across diverse tasks with improved cost efficiency.

    GPT-5.4 Nano

    GPT-5.4 nano is the most lightweight and cost-efficient variant of the GPT-5.4 family, optimized for speed-critical and high-volume workloads. It supports text and image inputs and is designed for low-latency tasks such as classification, data extraction, ranking, and sub-agent execution. Prioritizing responsiveness and efficiency over deep reasoning, GPT-5.4 nano is ideal for real-time systems, background processing, and distributed agent pipelines where minimizing cost and latency is essential.

    Happy Building.

  80. v2.2.28Feature

    System Update

    New Brand Identity

    New Logo

    Our new "Stacked A" mark features a dual-layer geometric design with metallic teal gradients, replacing the previous rainbow-arc logo. The layered depth effect represents the multiple AI providers unified behind a single API.

    What Changed
    • Logo: Geometric Stacked A with Apertis Teal (#2dd4bf → #0d9488)
    • Favicon: Full icon set across 25 sizes (16px–512px) for crisp rendering on every device - OG Images: All social preview images updated with new branding
    • Loading Animation: Inline SVG with breathing pulse effect replaces static image spinner
    • Header & Footer: Transparent logo mark with brand name for both light and dark modes
    What Didn't Change

    Your API keys, endpoints, SDK integrations, and billing — everything works exactly as before. This update is purely visual.

    Happy Building.

  81. v2.2.27Feature

    Feature Added

    ✨ Billing Credits API — Check Your Balance Programmatically

    We've launched a new API endpoint that lets you query your remaining credits and subscription quota using your API key — no dashboard login required.

    Endpoint: GET /v1/dashboard/billing/credits


    The Problem

    Until now, checking your Apertis balance meant opening the dashboard in a browser. This creates friction in several real-world scenarios:

    • Coding agents running overnight — Claude Code, Cursor, or Kilo Code sessions can burn through credits while you sleep. By the time you notice, the session has already failed mid-task with an insufficient balance error.
    • Team automation pipelines — CI/CD workflows that call AI APIs have no way to pre-check if there's enough budget before kicking off an expensive batch job.
    • Multi-key management — If you distribute API keys across projects or team members, there's no programmatic way to monitor which keys are running low.
    • Subscription cycle awareness — Subscription users couldn't check how much cycle quota remains without visiting the dashboard. Easy to accidentally exhaust your monthly allocation without realizing it.

    We looked at what other providers offer: OpenAI has no balance endpoint (this is one of the most requested features on their community forum). Anthropic's Admin API can query cost reports but not remaining credits, and requires a separate admin key. Neither provides a simple "how much do I have left?" API call.

    We decided to solve this properly.


    What It Returns

    A single request gives you the complete picture:

    PAYG users get their credit balance in USD:

    {
      "object": "billing_credits",
      "is_subscriber": false,
      "payg": {
        "remaining_usd": 12.50,
        "used_usd": 7.50,
        "total_usd": 20.00,
        "is_unlimited": false,
        "monthly_limit_usd": 50.00,
        "monthly_used_usd": 7.50,
        "monthly_reset_day": 1
      }
    }
    
    Subscription users see both their cycle quota and PAYG balance:
    
    {
      "object": "billing_credits",
      "is_subscriber": true,
      "payg": {
        "remaining_usd": 0.95,
        "used_usd": 0.05,
        "total_usd": 1.00,
        "is_unlimited": false
      },
      "subscription": {
        "plan_type": "pro",
        "status": "active",
        "cycle_quota_limit": 1000,
        "cycle_quota_used": 350,
        "cycle_quota_remaining": 650,
        "cycle_start": "2026-03-16T10:02:35Z",
        "cycle_end": "2026-04-16T10:02:35Z",
        "payg_fallback_enabled": true,
        "payg_spent_usd": 2.50,
        "payg_limit_usd": 10.00
      }
    }
    
    

    Use Cases
    1. Pre-flight budget check before expensive operations

    Before kicking off a large batch job or a long coding agent session, check if you have enough credits:

    import requests
    
    credits = requests.get(
         "https://api.apertis.ai/v1/dashboard/billing/credits",
          headers={"Authorization": "Bearer sk-your-key"}
      ).json()
    
    if credits["is_subscriber"]:
        remaining = credits["subscription"]["cycle_quota_remaining"]
        if remaining < 100:
            print(f"Warning: only {remaining} quota remaining in this cycle")
        else:
            remaining = credits["payg"]["remaining_usd"]
            if remaining < 1.0:
                print(f"Warning: only ${remaining:.2f} credits left")
    
    1. Automated low-balance alerts

    Set up a cron job or monitoring script that pings you when credits drop below a threshold:

      #!/bin/bash
      BALANCE=$(curl -s https://api.apertis.ai/v1/dashboard/billing/credits \
        -H "Authorization: Bearer $APERTIS_KEY" | jq '.payg.remaining_usd')
    
      if (( $(echo "$BALANCE < 5.0" | bc -l) )); then
        echo "Low balance alert: $BALANCE USD remaining" | \
          mail -s "Apertis Low Balance" [email protected]
      fi
    
    What Makes This Different

    This is an Apertis exclusive. We surveyed every major AI API provider:

    • OpenAI: No balance endpoint. The most upvoted feature request on their developer forum for over two years. Their Usage API shows historical spending but not remaining credits.
    • Anthropic: Admin API provides cost reports, but requires a separate admin key and doesn't return remaining balance.
    • Together AI, OpenRouter: No programmatic balance check.

    We believe knowing your balance should be as simple as making one API call. No special keys, no dashboard login, no scraping.

    Full API documentation

    Enjoy it.

  82. v2.2.26Feature

    Models Added

    Add GLM 5 Turbo

    GLM 5 Turbo

    GLM-5 Turbo is a high-performance model from Z.ai optimized for fast inference and agent-driven workflows. Designed for real-world environments such as OpenClaw scenarios, it delivers strong performance across long execution chains and complex task pipelines. The model features improved instruction decomposition, tool integration, scheduled and persistent execution, and enhanced stability for extended multi-step tasks, making it well suited for autonomous agents and production automation workflows.

    Enjoy it!

  83. v2.2.25Feature

    Models Added

    Add OpenRouter and NVIDIA models

    Nemotron 3 Super (Free)

    NVIDIA Nemotron 3 Super is a 120B-parameter open hybrid Mixture-of-Experts model designed for complex multi-agent and long-horizon reasoning workflows. It activates only 12B parameters per token, enabling high compute efficiency while maintaining strong accuracy on advanced tasks. Built on a hybrid Mamba–Transformer MoE architecture with multi-token prediction (MTP), the model delivers significantly higher token generation throughput than leading open models.

    Healer Alpha

    Healer Alpha is a frontier omni-modal model that integrates vision, audio understanding, reasoning, and action capabilities within a single system. It can natively perceive visual and auditory inputs, reason across multiple modalities, and execute complex multi-step tasks with precision, enabling advanced real-world agentic applications. Note: Prompts and completions processed by this model are logged by the provider and may be used for model improvement.

    Hunter Alpha

    Hunter Alpha is a frontier intelligence model with over 1 trillion parameters and a 1M-token context window, designed specifically for agentic applications. It excels at long-horizon planning, complex reasoning, and sustained multi-step task execution, delivering strong reliability and precise instruction following for advanced agent frameworks such as OpenClaw. Note: Prompts and completions processed by this model are logged by the provider and may be used for model improvement.

    These models are in free tier.

    Enjoy it!

  84. v2.2.24Feature

    Models Added

    Add Qwen3.5-9B & Seed-2.0-Lite

    Qwen3.5-9B

    Qwen3.5-9B is a multimodal foundation model from the Qwen3.5 family, built to deliver strong reasoning, coding, and visual understanding within an efficient 9B-parameter architecture. It adopts a unified vision-language design with early fusion of multimodal tokens, enabling the model to process and reason across text and images within the same context.

    With balanced multimodal capability and efficient deployment requirements, Qwen3.5-9B is well suited for applications that combine visual analysis, coding assistance, and general reasoning.

    Seed-2.0-Lite

    Seed-2.0-Lite is a balanced model designed for high-frequency enterprise workloads, optimizing for both capability and cost efficiency. It surpasses the previous-generation Seed-1.8 in overall performance while maintaining stable, production-ready quality. The model supports long-context processing, multi-source information fusion, multi-step instruction execution, and high-fidelity structured outputs.

    It is well suited for enterprise scenarios such as unstructured data processing, content generation, search and recommendation, and data analysis, delivering reliable results while significantly reducing operational cost.

    Enjoy it!

  85. v2.2.23Models

    Price Updated

    Update prices with 20% off for GPT-5.4 Models

    The prices for GPT-5.4 and GPT-5.4 Pro are updated to 20% off.

    Enjoy it.

  86. v2.2.22Feature

    Models Added

    Add GPT-5.4 and GPT-5.4 Pro

    GPT-5.4

    GPT-5.4 is OpenAI's latest frontier model, unifying the GPT and Codex lines into a single system designed for both general intelligence and advanced software engineering workflows. It supports text and image inputs and features a 1M+ token context window (≈922K input, 128K output), enabling high-context reasoning, coding, and multimodal analysis within a single workflow.

    The model delivers improved performance in coding, document understanding, tool use, and instruction following, and is designed as a strong default for complex tasks. It can generate production-quality code, synthesize information across large datasets, and execute multi-step workflows with fewer iterations and greater token efficiency.

    GPT-5.4 Pro

    GPT-5.4 Pro is OpenAI's most advanced model, built on the unified GPT-5.4 architecture with enhanced reasoning capabilities for complex and high-stakes tasks. It supports text and image inputs and features a 1M+ token context window (≈922K input, 128K output) for handling large-scale workflows and long-context analysis.

    Optimized for step-by-step reasoning, instruction following, and accuracy, GPT-5.4 Pro excels in agentic coding, long-context problem solving, and complex multi-step workflows, making it well suited for advanced engineering, research, and high-reliability applications.

    The discount prices will be updated in few days, stay tuned!

    Enjoy it.

  87. v2.2.21Feature

    Models Added

    ✨ New Models Added to All Subscription Plans — Free & Unlimited

    Hi there,

    Great news! We've just added three new models to every Apertis subscription plan at no extra cost — with unlimited usage.

    | Model | Highlights | |-------|-----------| | GLM 4.7 Flash | Zhipu's latest high-speed reasoning model — low latency, high throughput | | GPT-5.1 Codex (Mini) | OpenAI's lightweight code generation model — ideal for everyday dev tasks | | MiniMax M2.1 | MiniMax's next-gen general-purpose model — strong multilingual capabilities |

    How to Get Started

    These models are available immediately for all subscribers — no configuration changes needed. Simply use the model ID in your API calls:

    model: "glm-4.7-flash" model: "gpt-5.1-codex-mini" model: "minimax-m2.1"

    Why Free & Unlimited?

    We continuously partner with leading AI providers to bring high-value models into your subscription. More models, same price — that's our commitment to maximizing the value of your plan.

    Not a subscriber yet? Explore our plans →

    Questions? Feel free to reach out anytime.

    Happy Building.

  88. v2.2.20Feature

    Feature Added

    ✨ Apertis Python SDK v0.2.1 — Context Compression Now Available

    Hi there,

    We're excited to announce Apertis Python SDK v0.2.1, now with Context Compression support.

    What's New

    Context Compression automatically compresses prompt tokens in long conversations and large context scenarios — reducing API costs while maintaining response quality.

    Quick Start
      pip install apertis==0.2.1
    
    Resources
    • PyPI: https://pypi.org/project/apertis/0.2.1/
    • GitHub Release: https://github.com/apertis-ai/python-sdk/releases/tag/v0.2.1

    See the Release Notes for detailed usage and configuration options.

    Questions? Feel free to reach out anytime.

    Happy Building.

  89. v2.2.19Feature

    Models Added

    Add Gemini 3.1 Flash Lite Preview & GPT-5.3 Chat

    • GPT-5.3 Chat
    • Gemini 3.1 Flash Lite Preview
  90. v2.2.18Feature

    Feature Added

    ✨ New Feature: Context Compression

    Full notes on apertis.ai

  91. v2.2.17Feature

    Models Added

    Add Grok 4.2

    Grok 4.2 is the next major iteration of xAI's Grok series, advancing the model's reasoning, coding, and multimodal capabilities with architectural improvements over Grok 4 and 4.1. It is positioned as a more powerful and general-purpose frontier AI model in the Grok family with stronger deep reasoning and real-world task performance.

  92. v2.2.16Feature

    Models Added

    Add Qwen 3.5 Full Series & Seed-2.0-Mini

    The full Qwen 3.5 series is provided at Apertis Coding Plan as well, Enjoy it.

  93. v2.2.15Feature

    Models Added

    Add Nano Banana 2 (Gemini 3.1 Flash Image Preview)

    Gemini 3.1 Flash Image Preview (also known as "Nano Banana 2") is Google's latest state-of-the-art image generation and editing model, delivering Pro-level visual quality at Flash-level speed. It combines strong contextual understanding with fast, cost-efficient inference, enabling high-quality image generation and seamless iterative editing. Optimized for both performance and accessibility, it makes advanced visual creation workflows faster and more scalable.

  94. v2.2.14Feature

    System Update

    Cached responses now support streaming (SSE) delivery, covering ~80% of API traffic that uses stream: true.

    New feature: Cached responses now support streaming (SSE) delivery, covering ~80% of API traffic that uses stream: true.

    • On cache hit, the system emits synthetic SSE chunks from the stored response — no upstream API call needed
    • Content is split on rune boundaries (50 runes/chunk, 10ms intervals) to preserve multi-byte characters
    • Proper X-Cache-Hit, X-Cached-Tokens, and X-Actual-Model headers on streaming cache hits
    • Non-streaming cache hits continue to work as before (direct JSON response)

    Cache Correctness Hardening

    • Temperature guard: Only caches requests where temperature: 0 is explicitly present in the raw JSON body. Omitted temperature (Go zero value 0.0) is no longer falsely treated as cacheable — providers default to ~1.0 for omitted values
    • SSE error safety: If synthetic SSE emission fails mid-stream, the handler returns immediately instead of falling through to normal processing, preventing HTTP double-write corruption
    • Tool call exclusion: Responses containing tool_calls are excluded from cache storage since the SSE emitter only supports text content replay

    Cache TTL & Infrastructure

    • Default prompt cache TTL extended from 10 → 30 minutes

    Enjoy it.

  95. v2.2.13Feature

    Feature Added

    ✨ New Feature: Monthly Budget Controls

    You can now set a monthly spending cap on your API usage. Once enabled, usage is tracked against your limit and automatically resets on your chosen billing cycle date.

    • Monthly spending limit — Set a dollar amount that caps total API usage across all your keys each month.
    • Custom reset day — Choose any day from the 1st to the 28th as your monthly cycle start date.
    • Per-key budgets — Optionally allocate portions of your monthly limit to individual API keys for finer-grained control.
    • Usage alerts — Get an email notification when your usage crosses a customizable threshold (default 80%).
  96. v2.2.12Feature

    Models Added

    Add MiniMax M2.5 (Lightning)

    MiniMax-M2.5-Lightning is the high-speed variant of the M2.5 series, optimized for low latency, real-time responsiveness, and high-frequency workloads. It retains the core planning and execution strengths of M2.5 while further improving inference efficiency and response speed, making it ideal for interactive applications, rapid coding assistance, and workflow automation. With enhanced cost efficiency and reduced latency, M2.5-Lightning is particularly well suited for high-throughput, always-on deployments and production environments where speed and scalability are critical.

    https://apertis.ai/models/minimax-m2.5-lightning

  97. v2.2.11Feature

    Models Added

    Add Gemini 3.1 Pro Preview

    • Gemini 3.1 Pro Preview
  98. v2.2.10Feature

    Models Added

    Add Claude Sonnet 4.6

    • Claude Sonnet 4.6
  99. v2.2.9Feature

    Models Added

    Add Qwen3.5 397B A17B & Qwen3.5 Plus 2026-02-15

    • Qwen3.5 397B A17B
    • Qwen3.5 Plus 2026-02-15
  100. v2.2.8Feature

    Feature Added

    🚀 Apertis is now an Official Model Provider for Kilo Code

    Full notes on apertis.ai

  101. v2.2.7Feature

    Models Added

    Add MiniMax M2.5

    • MiniMax M2.5
  102. v2.2.6Feature

    Models Added

    Add GLM 5

    • GLM 5
  103. v2.2.5Feature

    Models Added

    Add Aurora Alpha, Qwen3 Max Thinking and Gemini Embedding 001

    • Aurora Alpha by OpenRouter
    • Qwen3 Max Thinking by Alibaba
    • Gemini Embedding 001 by Google
  104. v2.2.4Feature

    Models Added

    Add Pony Alpha

    • Pony Alpha
  105. v2.2.2Feature

    Models Added

    Add Claude Opus 4.6

    • Claude Opus 4.6
    • Claude Opus 4.6 (Thinking)
  106. v2.2.3Feature

    Models Added

    Add GPT-5.3-Codex

    • GPT-5.3-Codex
    • GPT-5.3-Codex (xhigh)
    • GPT-5.3-Codex (High)
    • GPT-5.3-Codex (Medium)
    • GPT-5.3-Codex (Low)
  107. v2.2.1Feature

    Feature Added

    ✨ Model Comparison is released

    From now on, at /models page and each model detail page, users are allowed to compare the costs and model functionality details at once within 3 models at most.

    Enjoy.

  108. v2.2.0Feature

    Feature Added

    🎉 Apertis Coding Plan Launched

    Full notes on apertis.ai

  109. v2.1.12Feature

    Models Added

    Add Qwen3 Coder Next

    • Qwen3 Coder Next
  110. v2.1.11Models

    Models Deprecated

    8 Models Deprecated

    • grok-4-fast:free
    • grok-4.1-fast:free
    • gpt-oss-120b:free
    • gpt-oss-20b:free
    • kimi-k2:free
    • gemini-2.5-pro-exp-03-25:free
    • minimax-m2:free
    • qwen3-235b-a22b-07-25:free
  111. v2.1.10Feature

    Models Added

    Add Kimi K2.5 and Qwen 3 Max (0123)

    • Kimi K2.5
    • Qwen 3 Max (0123)
  112. v2.1.9Feature

    Feature Added

    ✨ Ask AI for Documentation

    What's new
    • AI-powered documentation search
    • Ask questions in natural language and get precise, context-aware answers from Apertis docs.
    • No more manual searching or jumping between pages.
    • Built for integration
    • The Ask AI experience is powered by the same APIs you can use in your own product.
    • Easily embed an “Ask AI” assistant into your docs site, dashboard, or developer tools.
    • Developer-friendly by design
    • Optimized for technical questions, code usage, and API workflows.
    • Responses are grounded in official documentation to reduce hallucinations.
    • Faster onboarding & troubleshooting
    • Help users discover features, understand APIs, and resolve issues instantly.
    • Ideal for reducing support load and improving developer experience.
    Why it matters

    Ask AI turns documentation from a static reference into an interactive developer assistant. Whether you're building an internal tool, a public SDK site, or a customer-facing app, you can now offer the same AI-powered doc experience with minimal integration effort.

    Get started
    • Try Ask AI directly in Apertis Documentation
    • Follow the integration guide to add Ask AI to your own application
  113. v2.1.8Fix

    System Update

    Optimization on API requests: 40-80% latency reduction

    • Before: 30-55ms
    • After: 5-10ms (cache hits) / 15-25ms (cache misses)
    • Estimated improvement: 40-80% latency reduction
  114. v2.1.7Fix

    Feature Added

    Python SDK: Add comprehensive SDK features (v0.2.0)

    Add comprehensive SDK features (v0.2.0) New Features:

    • Vision/Image support with create_with_image() convenience method
    • Audio input/output support
    • Video content support
    • Web Search with create_with_web_search() convenience method
    • Reasoning Mode support (glm-4.7)
    • Extended Thinking (Gemini) support
    • Models API (list, retrieve)
    • Responses API (OpenAI format)
    • Messages API (Anthropic format)
    • Rerank API (BAAI/bge-reranker-v2-m3, Qwen models)
    • Helper utilities for base64 encoding

    Includes 81 unit tests covering all new functionality.

    PyPI Package: https://pypi.org/project/apertis/

  115. v2.1.6Fix

    Performance Improvement

    Python SDK - Fix API response compatibility issues (v0.1.1)

    Fix API response compatibility issues (v0.1.1)
    • Make id, object, created fields optional in ChatCompletion
    • Make id, object, created, model fields optional in ChatCompletionChunk
    • Make Usage fields optional to handle empty usage objects in streaming
    • Fix streaming by using httpx send() with stream=True instead of stream context manager

    Tested with real API calls:

    • Chat completions: working
    • Streaming: working
    • Embeddings: working
    • Tool calling: working
  116. v2.1.5Feature

    Feature Added

    🎉 Apertis Python SDK Released

    Full notes on apertis.ai

  117. v2.1.4Feature

    Feature Added

    🎉 Apertis Is Now an Official Community Provider in Vercel AI SDK

    • Great news — Apertis is now officially listed as a Community Provider in the Vercel AI SDK.

    Full notes on apertis.ai

  118. v2.1.3Feature

    Models Added

    Add GLM 4.7 Flash

    • GLM 4.7 Flash
  119. v2.1.2Feature

    Feature Added

    🎉 LiteLLM: Apertis is Now an Official LiteLLM Model Provider

    Full notes on apertis.ai

  120. v2.1.1Feature

    Feature Added

    🔥 LlamaIndex: Apertis is Now an Official LlamaIndex Model Provider

    Full notes on apertis.ai

  121. v2.1.0Feature

    Feature Added

    🎉 Release @apertis/ai-sdk-provider v1.1.1

    Full notes on apertis.ai

  122. v2.0.55Feature

    System Update

    Add Apple OAuth SSO

    We're excited to announce that Apple OAuth (Sign in with Apple) is now supported on the Apertis login page for all users.

    You can now sign in to Apertis using your Apple ID for a faster, more secure, and privacy-focused authentication experience.

    Why this matters:

    • ✅ One-click login with your Apple ID
    • 🔒 Enhanced security with Apple's OAuth flow
    • 🕵️ Privacy-first: hide your email if you choose
    • 🚀 Faster onboarding for new users

    This update is available to all users starting today.

    👉 Try it now on the login page and let us know what you think.

  123. v2.0.54Feature

    Feature Added

    New Feature: FREE Real-Time Web Search for Any Model

    We're excited to announce Web Search - a powerful new feature that brings real-time internet data to any AI model on Apertis.

    What's New

    Simply add :web to any model name to enable real-time web search:

    • gpt-4o-mini:web
    • claude-3-5-sonnet-20241022:web
    • gemini-1.5-pro:web,...

    🎉 Simply add :web suffix directly for ALL Model Ids!!!

    Key Features
    • Real-Time Data: Get up-to-date information on stock prices, news, weather, and more
    • Source Citations: Every response includes web_sources with referenced URLs
    • Streaming Support: See the 🔍 Web searching... indicator followed by live content streaming
    • Zero Extra Cost: Web search costs are absorbed by the platform - you only pay for model tokens
    • Graceful Fallback: If search fails, the request automatically continues without interruption
    🔥 Key benefits
    • Real-time information (stock prices, news, weather, live events)
    • Source citations included in every response
    • Streaming support with visible search indicators
    • 🤩 No extra cost — web search is included
    • Automatic fallback if search is unavailable to ensure stability

    Example Response

      {
        "content": "As of January 6, 2026, Apple (AAPL) is trading at $262.36...",
        "web_sources": [
          {"title": "Apple Investor Relations", "url": "https://investor.apple.com/..."},
          {"title": "MarketWatch", "url": "https://www.marketwatch.com/..."}
        ]
      }
    
    Documentation

    For complete usage details, request parameters, and best practices, check out our documentation:

    👉 https://docs.apertis.ai/web-search

  124. v2.0.53Feature

    Models Added

    Add Allen AI Models

    • Olmo 3.1 32B Instruct
    • Molmo2 8B (Free)
  125. v2.0.52Feature

    Models Added

    Add ByteDance Seed Models

    • Seed 1.6
    • Seed 1.6 Flash
  126. v2.0.51Feature

    Models Added

    Add GLM 4.7 & MiniMax M2.1

    • GLM 4.7
    • MiniMax M2.1
  127. v2.0.50Feature

    Models Added

    Add Gemini 3 Flash & GPT Image 1.5

    • Gemini 3 Flash Preview
    • GPT Image 1.5
  128. v2.0.49Feature

    Models Added

    Add Ai2 & Mistral AI models

    • Olmo 3.1 32B Think (Free)
    • Mistral Small Creative
  129. v2.0.48Feature

    Models Added

    Add Nemotron 3 Nano 30B A3B (Free)

    • Nemotron 3 Nano 30B A3B (Free)
  130. v2.0.47Feature

    Models Added

    Add GPT-5.2 Models

    • GPT-5.2
    • GPT-5.2 Pro
    • GPT-5.2 Chat
  131. v2.0.46Feature

    Models Added

    Add Devstral 2 2512

    • Devstral 2 2512
  132. v2.0.45Feature

    Models Added

    Add GLM 4.6V & Devstral 2 2512 (Free)

    • GLM 4.6V
    • Devstral 2 2512 (Free)
  133. v2.0.44Feature

    Models Added

    Add GPT-5.1-Codex-Max

    • GPT-5.1-Codex-Max
  134. v2.0.43Feature

    Models Added

    Add Ministral 3 Models

    • Ministral 3 14B 2512
    • Ministral 3 8B 2512
    • Ministral 3 3B 2512
  135. v2.0.42System

    Feature Added

    Release: New API Endpoints

    ✨ Overview

    We now supports the API schemas from both OpenAI and Anthropic, introducing two new unified endpoints:

    • /v1/responses – OpenAI Responses API compatible
    • /v1/messages – Anthropic Claude Messages API compatible

    This upgrade enables seamless multimodal input, structured outputs, and tool-use across all 410+ Apertis models, strengthening our position as a unified and future-ready AI gateway.


    🆕 Added

    1. /v1/responses – OpenAI Responses API Compatibility
    • Fully supports the official OpenAI Responses API specification
    • Backward-compatible with existing Chat Completions requests
    • Supports the flexible input[] content structure (text, image, multimodal blocks)
    • Supports response_format including strict JSON Schema, tool calling, and structured outputs
    • Streaming responses use the new OpenAI-aligned SSE protocol
    • Enabled for all 410+ Apertis models, including reasoning and open-source models

    2. /v1/messages – Anthropic Claude Messages API Compatibility
    • Full compatibility with Anthropic’s Claude Messages API
    • Supports multi-part content blocks (text / image / tool)
    • Supports strict JSON Schema (response_format: json_schema)
    • Enables Claude-series features such as tool use, vision, and system prompts
    • Functions as a drop-in replacement for the official Claude SDK
    • Automatically normalizes and harmonizes content block structures for cross-provider compatibility
  136. v2.0.41Feature

    Models Added

    Add Mistral and Amazon Models

    • Mistral Large 3 2512
    • Nova 2 Lite
    • Nova 2 Lite (Free)
  137. v2.0.40Feature

    Models Added

    Add DeepSeek v3.2 models

    • DeepSeek V3.2
    • DeepSeek V3.2 Speciale
  138. v2.0.39Feature

    Models Added

    Add Grok 4.1 Series Models

    • Grok 4.1 Fast
    • Grok 4.1 Fast (Free)
    • Grok 4.1 (Thinking)
    • Grok 4 Image
  139. v2.0.38Feature

    Models Added

    Add TNG Tech Models

    • R1T Chimera (Free)
    • R1T Chimera
    • DeepSeek R1T2 Chimera
    • DeepSeek R1T Chimera
  140. v2.0.37Feature

    Models Added

    Add Olmo 3 Models

    • Olmo 3 7B Think
    • Olmo 3 7B Instruct
    • Olmo 3 32B Think
  141. v2.0.36Feature

    Models Added

    Add Claude Opus 4.5

    • Claude Opus 4.5
  142. v2.0.35Feature

    Models Added

    Add Gemini 3 Pro Image Preview

    • Gemini 3 Pro Image Preview
  143. v2.0.34Feature

    Models Added

    Add Grok 4.1 Fast (Free)

    • Grok 4.1 Fast (Free)
  144. v2.0.33Feature

    Models Added

    Add Gemini 3 Pro Series

    • Gemini 3 Pro Preview
    • Gemini 3 Pro Preview (Thinking)
  145. v2.0.32Feature

    Models Added

    Add GTP-5.1 Codex Series

    • GPT-5.1 Codex
    • GPT-5.1 Codex (High)
    • GPT-5.1 Codex (Medium)
  146. v2.0.31Feature

    Models Added

    Add GPT-5.1 Series Models

    • GPT-5.1
    • GPT-5.1 (Thinking)
    • GPT-5.1 (High)
    • GPT-5.1 (Medium)
    • GPT-5.1 (Low)
  147. v2.0.30Models

    Price Updated

    Sora 2 Pro

    • Updated Price: $2.5/per use
  148. v2.0.29Feature

    Models Added

    Add Gemini Pro 3

    • Gemini 3 Pro
    • Gemini 3 Pro (Thinking)
  149. v2.0.28Feature

    Models Added

    Add 3 Models

    • Kimi K2 Thinking
    • Polaris Alpha
    • Kimi Linear 48B A3B Instruct
  150. v2.0.27Feature

    Models Added

    Add 4 Models

    • Sonar Pro Search
    • Codestral Embed 2505
    • Mistral Embed 2312
    • Nova Premier 1.0
  151. v2.0.26Feature

    Models Added

    Add 5 Models

    • DeepSeek-OCR
    • MiniMax M2
    • Nemotron Nano 12B 2 VL
    • Nemotron Nano 12B 2 VL (Free)
    • GPT OSS Safeguard 20B
  152. v2.0.24Feature

    Models Added

    Add Qwen3 VL 32B Instruct

    • Qwen3 VL 32B Instruct
  153. v2.0.23Feature

    Models Added

    Add MiniMax M2 (Free)

    • MiniMax M2 (Free)
  154. v2.0.22Feature

    Models Added

    Add Andromeda Alpha

    • Andromeda Alpha
  155. v2.0.21Feature

    Price Updated

    Price Update on Claude Haiku 4.5 with Less Cost

    • Claude Haiku 4.5 Input: 0.0008/per k tokens, Output: 0.004/per k tokens
  156. v2.0.20Feature

    Models Added

    Add Sora 2 Models: Sora 2 Pro, Sora 2 HD, Sora 2, Sora 2 Landscape, Sora 2 Portrait

    • Sora 2 Pro
    • Sora 2 HD
    • Sora 2
    • Sora 2 Landscape
    • Sora 2 Portrait
  157. v2.0.19Feature

    Models Added

    Add Claude Haiku 4.5 & Claude Haiku 4.5 (Thinking)

    • Claude Haiku 4.5
    • Claude Haiku 4.5 (Thinking)
  158. v2.0.18Feature

    Models Added

    Add o3 Deep Research & o4 Mini Deep Research

    • o3 Deep Research
    • o4 Mini Deep Research
  159. v2.0.17Feature

    Models Added

    Add Qwen3 VL 8B Instruct & Qwen3 VL 8B Thinking

    • Qwen3 VL 8B Instruct
    • Qwen3 VL 8B Thinking
  160. v2.0.16Feature

    Models Added

    Add GPT-5 Pro

    • GPT-5 Pro
  161. v2.0.15Feature

    Models Deprecated

    Deprecate OpenRouter Models

    • Sonoma Sky Alpha
    • Sonoma Dusk Alpha
  162. v2.0.14Feature

    Models Added

    Add Llama 3.3 Nemotron Super 49B V1.5 & ERNIE 4.5 21B A3B Thinking

    • Llama 3.3 Nemotron Super 49B V1.5
    • ERNIE 4.5 21B A3B Thinking
  163. v2.0.13Feature

    Models Added

    Add Gemini 2.5 Flash Image

    • Gemini 2.5 Flash Image
  164. v2.0.12Fix

    Performance Improvement

    Update API Docs for LangChain

  165. v2.0.11Feature

    Feature Added

    Latest Docs on Context7 MCP Server

    Now, you can access Stima API Docs with Context7 MCP Server during development journey, enjoy it!!!

  166. v2.0.10Feature

    Models Added

    Add 2 Models

    • GLM 4.6 (Thinking)
    • Qwen3 Coder Flash
  167. v2.0.9Models

    Price Updated

    Update Model Price: GLM 4.6

  168. v2.0.8Models

    Price Updated

    Update Model Price: DeepSeek V3.2 Exp

  169. v2.0.7Feature

    Models Added

    Add 1 Model

    • GLM 4.6
  170. v2.0.6Feature

    Models Added

    Add 2 Models

    • Claude Sonnet 4.5
    • DeepSeek V3.2 Exp
  171. v2.0.5Feature

    System Update

    Add Stripe Payment History (hosted by Stripe) to Topup Page

  172. v2.0.4Feature

    Models Added

    Add 3 Models

    • Gemini 2.5 Flash Lite Preview 09-2025
    • Gemini 2.5 Flash Preview 09-2025
    • Grok 4 Fast
  173. v2.0.3Feature

    Models Added

    Add 1 Model

    • Grok 4 Fast (Free)
  174. v2.0.2Feature

    Models Added

    Add 4 Models

    • Qwen3 Coder Plus
    • Qwen3 Max
    • Qwen3 VL 235B A22B Instruct
    • Qwen3 VL 235B A22B Thinking
  175. v2.0.1System

    System Update

    Update Changelog Articles Management

  176. v2.0.0System

    System Update

    Major System Upgrade: Database Migration to Supabase

    image

    We are pleased to announce that Stima API has successfully completed a major infrastructure upgrade by migrating our entire database system to Supabase. This migration represents a significant milestone in our commitment to data security, compliance, and system reliability.

    🔒 Why Supabase?

    1. Enterprise-Grade Security Compliance
    • ✅ SOC 2 Type II Certified - Ensuring highest level of security controls
    • ✅ ISO 27001 Certified - Meeting international information security standards
    • ✅ HIPAA Compliant - Ready for sensitive data processing requirements
    2. Enhanced Data Protection
    • End-to-end encryption in transit
    • Automatic backup and disaster recovery
    • Encrypted data storage (AES-256)
    • Granular access control
    3. Improved System Performance
    • 99.9% uptime SLA guarantee
    • Global distributed architecture for reduced latency
    • Auto-scaling capabilities for traffic spikes
    • Real-time data synchronization

    📈 Impact on Your Experience

    • Zero-downtime migration - Continuous service availability
    • Data integrity - All historical data preserved
    • Faster response times - Optimized database architecture
    • Enhanced security - Enterprise-grade protection for your data

    🚀 Looking Forward

    This migration enables us to:

    • Deliver more reliable AI model services
    • Implement stricter data privacy measures
    • Comply with global data protection regulations (GDPR, CCPA, etc.)
    • Build a solid foundation for future feature expansion

    Thank you for your continued trust and support. Should you have any questions, please don't hesitate to contact our technical support team.