AI Test Engineer

Beijing-based enterprise AI company building foundation-model, agentic-AI, and AI-transformation products, including the TrueNorth enterprise decision hub.

Beijing, China

Funding history

About 01.AI

01.AI develops full-stack enterprise AI solutions, industry-agent applications, and sovereign/industry model-training capabilities.

View jobs by 01.AI

Skills

Candidate Availability

Required and preferred rules are kept separate and reflect the wording in the original posting.

About the Role

You will own testing and evaluation delivery for enterprise Agent projects. You will define test strategies, evaluation plans, and acceptance criteria; identify quality risks; and provide launch, delivery, and customer-acceptance recommendations. You will evaluate AI applications, knowledge retrieval, Agents, tool calls, multi-turn conversations, and workflows. You will build enterprise evaluation systems, datasets, benchmarks, regression mechanisms, Harness capabilities, and evaluation-platform components. You will also lead quality initiatives for stability, performance, security, and adversarial testing.

Requirements

  • 5 or more years of testing, test development, quality, or evaluation experience
  • 2 or more years of AI or Agent evaluation experience
  • Ability to independently deliver complex projects
  • Harness knowledge and practical testing or evaluation Harness experience
  • AI and Agent evaluation methods
  • Evaluation dataset development, metric design, and combined human and automated evaluation
  • Knowledge retrieval, Agent, tool calling, multi-turn conversation, and workflow knowledge
  • Python
  • Enterprise project experience
  • Data isolation, permission management, auditing, release rollback, and private deployment quality requirements
  • Cross-team delivery ability

Responsibilities

  • Define test strategies, evaluation plans, and acceptance criteria for enterprise Agent projects
  • Identify quality risks from customer business and usage scenarios
  • Drive product, algorithm, engineering, and delivery teams to resolve issues
  • Deliver evaluation conclusions and improvement recommendations for launch, delivery, and customer acceptance
  • Test and evaluate AI applications, knowledge retrieval, Agents, tool calls, multi-turn conversations, and workflows
  • Design combined human, automated, and model-assisted evaluation approaches
  • Analyze logs and call chains, drive fixes, and verify regressions
  • Build quality metrics, evaluation datasets, standards, baselines, release thresholds, and feedback mechanisms
  • Build Harness and evaluation-platform capabilities for execution, scoring, comparison, analysis, and reporting
  • Lead stability, performance, security, and adversarial testing initiatives
  • Develop testing and evaluation scripts, tools, or platform components