Grok 4.5 Release: Engineering Performance and Agent Capabilities with SpaceXAI and Cursor
AI動向 業界ニュース 8 min read

Grok 4.5 Release: Engineering Performance and Agent Capabilities with SpaceXAI and Cursor

This post explores the release of Grok 4.5 by SpaceXAI and its integration with Cursor. With a 62.0% score on the DeepSWE 1.0 benchmark, the model demonstrates enhanced engineering capabilities and interdisciplinary reasoning. Notable features include high-efficiency performance for agentic tasks using "One Prompt," alongside cost-effective pricing. It offers developers a powerful new tool for streamlining complex workflows and coding tasks.

Grok 4.5 Released: SpaceXAI and Cursor's "Revenge" Counterattack—The Rules of the Code World Have Changed

If competition in the field of Large Language Models is a marathon without a finish line, then what SpaceXAI and Cursor have staged this past month is an astonishing sprint.

When I saw the news of the Grok 4.5 release, what struck me most wasn't just its terrifying capabilities, but the resolution and drive behind it—a "revenge" move after feeling toyed with by OpenAI and Anthropic. As engineers, we understand that feeling all too well: in this competitive arena, the only way to make the internal friction and grievances vanish is to win the battle through the "hard power" of product performance.

A "Performance Monster" in the Eyes of Engineers

The release of Grok 4.5 isn't just a simple model iteration, but a precision strike against complex engineering tasks. It hasn't just "become smarter"; it has genuinely learned how to think, debug, and build like an engineer.

Observing the DeepSWE 1.0 and SWE Bench Pro benchmarks, we can clearly see its positioning:

  • Leap in Engineering Capabilities: In the DeepSWE 1.0 benchmark, Grok 4.5 achieved an impressive score of 62.0%, ranking just behind Fable and GPT 5.5, proving its powerful potential in solving real-world engineering tasks.
  • More Than Just Code: Grok 4.5's training data is not limited to codebases but covers high-quality STEM tasks, research papers, and various types of knowledge-work data. This means it has a deeper foundation when handling logically complex tasks that require interdisciplinary reasoning.
  • Extreme Efficiency: This point impressed me the most. Grok 4.5 runs at 80 TPS (tokens per second), and in SWE Bench Pro tasks, its output token efficiency is 4.2 times that of Opus 4.8 (max). Lower token consumption means faster response speeds and lower costs.

"One Prompt": A Minimalist Path from Inspiration to Reality

What surprised me most about Grok 4.5 is its performance in "Agentic Tasks." In the demonstration, with just a simple Prompt, it was able to build a complex Solar System simulator or an exquisite Q3 business review presentation.

This "idea-to-product" capability greatly reduces the trivial burdens in our daily work:

  • Frontend & Visualization: Capable of directly using Three.js to generate a modern-designed space simulation HUD, complete with interactive timeline controls.
  • Office Automation: It is no longer just writing documents; it can directly build complex models involving multi-page formulas in Excel or automatically layout graphics in PPT.

Why Should We Pay Attention to This Release?

This is not just a model upgrade; it symbolizes a paradigm shift of a "geek counterattack." SpaceXAI has invested tens of thousands of NVIDIA GB300 GPUs and performed deep stability optimizations for large-scale operations. This control over the underlying infrastructure allows it to achieve highly asynchronous training, demonstrating high robustness when processing long-link agentic tasks.

Moreover, for us developers, its pricing strategy is highly competitive: $2 per million input tokens and $6 per million output tokens. Combined with its high token efficiency, this makes the deployment cost of large models in production environments much more controllable.

Conclusion: For That Grand Victory

Behind Grok 4.5 lies a team determined to prove themselves. In this ever-changing AI era, we are all looking for that "worldly success"—not to prove others wrong, but to prove that we are in the right lane and possess the ability to change the rules.

Now, Grok 4.5 is available via Cursor or API call. Whether you want to experience its code-building capabilities or test its agentic workflows, this is undoubtedly one of the most worthwhile technical tools to try right now.

Go give it a try; perhaps in the next Prompt, you too can feel the thrill of that "youthful spirit" returning.

VIBECODING

Readable articles from the intersection of AI and real-world development.

© 2026 VibeCoding Japan, Inc. All Rights Reserved.