Skip to main content
Togal.AI

AI QA Engineer / AI Test Architect

RemoteEMEA
Published
Role
QA
Employment
Full-time
Company size
Startup
Salary not disclosed
Check eligibility

Open to Anywhere in EMEA. Set where you work from to check your eligibility.

Core skills

k6/JMeterDeepEval

Required skills

Playwright/CypressTypeScript/PythonGitHub Actions/CircleCIPostmanDockerKubernetes

Optional skills

LangfuseRAGAS

What you'll do

  • Own quality strategy end-to-end: define, implement, and continuously evolve testing standards across functional, non-functional, and AI-specific dimensions, ensuring quality is embedded from requirements through production.
  • Build and maintain non-functional test automation: design and run performance, load, and stress test suites (k6, JMeter, Gatling etc.) integrated directly into CI/CD pipelines, with quality gates that protect every release.
  • Design and operate (or contribute to) LLM/AI eval frameworks: establish evaluation pipelines (using tools such as DeepEval, Langfuse etc.) to assess AI feature quality across metrics including accuracy, hallucination rate, relevance, faithfulness, and safety.
  • Test AI features and agentic behaviours: validate non-deterministic outputs, prompt variability, model regression, guardrail enforcement, and multi-step agent task-completion rates as first-class quality concerns.
  • Champion shift-left and continuous testing: embed QA into planning, design review, and sprint ceremonies so defects are caught before they're coded, not after they ship.

What they require

  • Traditional QA foundations: solid understanding of deterministic testing: test planning, test case design, functional/regression/exploratory testing, defect lifecycle management, and quality metrics.
  • Test automation engineering: deep expertise in writing and maintaining automated test suites using modern frameworks (Playwright, Cypress, or similar) with at least one modern programming language, such as TypeScript (strongly preferred) or Python, specifically for building robust test libraries.
  • Non-functional test automation: hands-on experience designing and running performance, load, and stress tests with tools such as k6 or JMeter, including CI/CD integration and threshold-based quality gates.
  • AI/LLM testing literacy: practical understanding of what makes AI systems non-deterministic, and experience (or strong working knowledge) of testing LLM-based features for hallucination, consistency, safety, and latency.
  • Eval framework awareness: a working understanding of LLM evaluation concepts: scoring metrics (BLEU, ROUGE), LLM-as-judge patterns, and familiarity with at least one eval framework (DeepEval, RAGAS, etc.).

Benefits

  • Join a dynamic team of AI-native engineering team.
  • AI-native from day one - you'll be building the quality discipline for a product that uses AI at its core, making every quality decision novel and impactful.
  • Be the quality voice, not a quality follower - this role has direct influence over how Togal defines and measures product excellence.
  • A culture of innovation, continuous learning, and high growth.
  • Comprehensive benefits

We are an innovative technology company providing a cutting-edge, AI-powered cloud platform for the construction industry. Created by industry experts with deep estimating experience, our software dramatically streamlines the pre-construction process. Our solution uses advanced machine learning to automate traditionally time-consuming takeoff tasks, helping estimators work up to 80% faster while reducing costly errors. Our collaborative platform enables real-time teamwork, instant drawing analysis, and features a revolutionary conversational AI interface that transforms how professionals interact with construction plans. Founded by construction industry veterans, our award-winning application automates the takeoff process, enabling estimators to analyze blueprints in seconds rather than hours or days.

ConstructionStartup
Salary not disclosed