QA / ML Tester needed to ensure quality of AI evaluation frameworks for enterprise platforms. Customer is a multinational corporation with offices in 180+ countries, focusing on reduced-risk products and IoT solutions.
Обов’язки
- Design, implement, and maintain automated test suites for AI evaluation frameworks and evaluation pipelines.
- Develop Python-based test automation using pytest to validate evaluator behavior, quality scoring, and framework reliability.
- Create and maintain known-good and known-bad test datasets, sessions, and workflows for evaluator correctness validation.
- Validate the accuracy and consistency of evaluation results across different agent workflows, prompts, tools, and execution scenarios.
- Design and execute integration tests for on-demand evaluation workflows integrated into CI/CD pipelines.
- Verify online evaluation behavior, sampling accuracy, and evaluation result consistency in production-like environments.
- Conduct functional testing of evaluation components, including evaluator execution flows, scoring logic, and result aggregation.
- Collaborate with AI and platform engineering teams to identify edge cases, failure scenarios, and evaluation blind spots.
- Validate workflow compliance, tool execution assessment, and end-to-end quality evaluation processes.
- Support feasibility assessments for applying evaluation frameworks to non-AgentCore runtimes and alternative AI execution environments.
- Analyze defects, inconsistencies, and quality regressions within evaluation systems and provide actionable recommendations.
- Contribute to quality assurance standards, testing methodologies, and best practices for AI evaluation platforms.
Вимоги
- Python test automation (pytest)
- Evaluator correctness testing (known-good / known-bad session pairs)
- On-demand mode integration testing with CI/CD
- Online mode sampling accuracy verification