量子位

Lenovo Tianxi’s self-developed Code AI TianxiCode wins the global first place at SWE-bench-Live

Lianxing TianxiAI’s independently developed professional code intelligence framework TianxiCode ranks first globally with a 71% problem-solving rate

Image source · 量子位

Recently, in the SWE-bench-Live (Lite ranking) software engineering evaluation test, which is internationally recognized as one of the most challenging and closely mimics real development environments, the professional code intelligence framework TianxiCode developed by Lenovo Tianxi AI, in conjunction with DeepSeek-v4.1-Flash, achieved breakthrough results: it ranked first globally with a 71% problem-solving rate and was officially approved by the official authorities.

This achievement not only marks Lenovo’s self-developed code intelligence framework as one of the top international players, but also demonstrates the hardcore strength of TianxiAI’s ecosystem in leveraging its self-developed engineering architecture to drive model potential and compete with world-class solutions in complex software engineering scenarios. Tianxi Code is a sub-capability of TianxiAI focused on code generation and engineering development, and it will be applied to Lenovo’s AI hardware products in the future.

What is the value of SWE-bench-Live in real service scenarios?

Unlike traditional code tests that focus on simple grammar or single-function generation, SWE-bench-Live is an authoritative dynamic benchmarking standard in the global software engineering field.

SWE-bench-Live is based on issues in real GitHub projects. It evaluates the ability of models and agents to generate code patches and fix problems, and provides an reproducible execution environment. Compared to simply testing code snippet generation, this type of evaluation gets closer to the problem-solving processes encountered in real development. It sets comprehensive requirements for understanding the system, identifying problems, and completing fixes.

To ensure absolute fairness in the evaluation, SWE-bench-Live has established a “Verified” verification mechanism. The evaluation proposals must submit the full chain of AI operation trajectories, and the official team will strictly review the evaluation inputs and information isolation environment, thoroughly checking for any possibility of leaking standard answers, test cases, or results to the AI. TianxiCode and DeepSeek-v4.1-Flash successfully passed the review and received the “Verified” certification, which fully proves the reproducibility and high quality of their code fixes.

TianxiCode architecture breakthrough, unleashing engineering-level implementation capabilities

If the entire “TianxiCode with DeepSeek-v4.1-Flash” system is compared to a software engineer, then DeepSeek-v4.1-Flash is the engineer’s brain, responsible for reading code, understanding semantics, and coming up with solutions; while TianxiCode is the engineer’s eyes, hands, toolbox, and work flow guidelines. No matter how intelligent the brain is, without a pair of hands capable of typing code and checking logs, as well as strict working habits, it cannot solve bugs on complex real-world projects alone.

Multiple list comparison data confirm this: without system collaboration based on a top-level intelligent agent architecture, even when using top-tier closed-source models, it often remains helpless in dealing with real engineering challenges. These facts prove the technical value of TianxiCode.

In real software engineering scenarios, TianxiCode demonstrates three major system engineering capabilities: cross-file multi-hop retrieval and precise context management, self-driven planning and multiple tool calls, as well as closed-loop patch generation and test-driven self-healing.

Multi-Hop Retrieval and Context Pruning & Slicing enable TianxiCode to find issues like a seasoned engineer would, quickly locating the root causes of defects in large code repositories.

Autonomous Planning and Multi-turn Tool Calling enable TianxiCode to “type while watching” like a real programmer: when a bug occurs, it first outlines a troubleshooting plan, actively opens the terminal command line to run tests, reviews logs, and compares modification records. If it encounters an unsolvable problem, it can also flexibly change its approach to continue troubleshooting, achieving autonomous progress.

Closed-Loop Patching and Test-Driven Self-Correction enable TianxiCode to self-test and self-correct. TianxiCode will automatically run the code in an isolated environment, and immediately reflect on and modify it once errors or damage to other functions are detected, until the patch passes the project review.

It can be said that TianxiCode’s solid self-developed engineering foundation demonstrates the decisive power of a robust intelligent agent framework in complex real-world software development scenarios.

Moving from the top of the list to the depths of the industry, Tianxi Ecology accelerates AI accessibility

As a key research and development position in the professional technology field within the TianxiAI ecosystem, TianxiCode has ranked among the top tier in international top-tier evaluations. This not only serves as a concentrated verification of its technical strength, but also paves the way for the implementation of smart agents in real-world development scenarios.

The ultimate value of the code agent lies not in a single ranking on a list, but in sharing the heavy costs of debugging and engineering collaboration with frontline developers. In the future, Lenovo TianxiAI will continue to transform TianxiCode’s technical achievements into practical toolchains for developers. With a mature autonomous agent architecture, professional code tools that are secure, efficient, and truly capable of solving problems independently will be deeply integrated into Lenovo’s hardware devices.

This article is provided by Lenovo, and Qbit has been authorized to reproduce it. The views belong to the original author.

Original source

量子位

Content notes

Original publication and rights belong to the source.

Machine translation · Refer to the original