QA Engineer — AI-augmented Test Automation

Full-time · Hybrid

← All open roles

QA Engineer — AI-augmented Test Automation

Full-time · Hybrid · Essential

A builder QA role: unit tests for critical business logic, Playwright and HTTP regression, AI test agents on every shipped ticket, and scheduled LLM quality evals.

The role

Essential curates small teams of world-class engineers to accelerate our clients' innovation cycles. For this role you will own product quality for a production AI companion: users talk to an AI avatar via video, voice, or text, powered by a self-hosted LLM stack, Supabase, Stripe billing, a React/TypeScript frontend, and an iOS app.

This is a builder QA role, not a manual-execution role — most execution is automated or delegated to AI agents, and you build that automation.

The one non-negotiable: you test with AI agents

Hands-on daily use of agents to write, run, and triage tests: generating specs from tickets, driving exploratory browser testing, delegating fixture creation — combined with a clear view of their failure modes (hallucinated passes, brittle selectors, scope drift). You know why a non-deterministic agent must never be the sole release gate.

What you will do

  • Close the business-logic gap: unit tests for entitlement and token math, per-handler tests for streaming edge functions with mocked LLM fixtures.
  • Run an AI feature-test agent on each shipped ticket: read the ticket and diff, produce a test plan, execute it via browser automation, and post evidence. Graduate repeatable checks into the permanent regression suite.
  • Put LLM quality monitoring on a schedule: judge-scored quality canaries and scores that make drift visible over weeks.
  • Maintain Playwright (login, text chat, spoken voice and video sessions) and HTTP regression (liveness, auth gates, safety and prompt-injection probes).
  • Be the quality voice in the ticket flow: review acceptance criteria, advise on promotion, and shrink the post-release manual checklist.

Must-have

  • 3+ years in QA / test automation, with real code ownership of test suites in TypeScript/JavaScript.
  • Playwright (or equivalent: Cypress, WebdriverIO) in CI, including flaky-test diagnosis and stabilisation.
  • Hands-on daily use of AI agents for testing — you will be asked to show this live.
  • Comfortable with API testing, SQL for read-only data verification, reading logs, and GitHub Actions.
  • Structured thinker who writes concise, evidence-backed reports that engineers act on without a follow-up call.
  • Fluent written English; German strongly preferred (DE is the primary UI locale).

Nice-to-have

  • Testing LLM-based products: eval design, judge models, rubric scoring, safety/red-team probing.
  • Testing real-time media (WebRTC audio/video, fake media devices), voice UX.
  • Vitest / Deno test, Supabase, Stripe test mode, iOS TestFlight / Capacitor, accessibility and i18n testing.
  • Experience in health or other safety-sensitive domains.

Working principles

  • Automation and agent leverage first, manual execution last. Manual work that recurs three times must become a script, a Playwright spec, or an agent skill.

How we assess

We assess this hands-on: given a ticket and a deployed URL, you produce a test plan, execute it with an agent driving the browser, and deliver a report — using your own tooling. We watch how you brief the agent, what you verify, and how you recover when it goes wrong.