Careers / Remote

Build the systems that make AI accountable.

Join a small, senior team turning production AI from guesswork into measurable engineering.

WORKING MODEL
01
Remote by defaultWork where you do your best thinking.
02
Async firstClear writing over calendar overload.
03
Evidence ledMeasure outcomes, not activity.
HIRING SYSTEM ONLINE

Open positions

Choose your problem space.

Every role is remote. Apply to the one where your experience is strongest.

OPENING / 01

AI Evaluation Engineer

Build the test systems that show whether production AI agents are getting better or quietly breaking.

Full-timeRemote

What you'll do

  • Turn real production cases into representative golden datasets.
  • Design deterministic checks and LLM judges, then validate them against human labels.
  • Wire evaluation suites into delivery pipelines so regressions block releases.
  • Analyze failure clusters and translate them into measurable engineering work.

What you bring

  • Strong Python and software testing fundamentals.
  • Hands-on experience evaluating LLM applications, agents, or retrieval systems.
  • Comfort with experiment design, metrics, and imperfect real-world data.
  • Clear written communication in an async remote environment.
Bonus

Experience with tracing platforms, prompt testing, or human annotation workflows.

PythonLLM evaluationCI/CD
Apply for this role
OPENING / 02

Applied AI Engineer

Ship production agent workflows whose reliability can be measured, explained, and improved.

Full-timeRemote

What you'll do

  • Build agentic workflows for support, document processing, and back-office operations.
  • Integrate model APIs, business tools, retrieval, and structured outputs.
  • Instrument every critical step for quality, latency, and cost.
  • Work with evaluation engineers to move workflows safely from prototype to production.

What you bring

  • Strong TypeScript and backend application development experience.
  • Experience shipping LLM features with tool use, structured outputs, or retrieval.
  • Sound judgment around retries, fallbacks, validation, and failure handling.
  • Ability to own a problem from discovery through production support.
Bonus

Python experience or familiarity with workflow orchestration and model routing.

TypeScriptAgentsTooling
Apply for this role
OPENING / 03

AI Observability Engineer

Make model behavior visible through traces, cost telemetry, scorecards, and useful operational signals.

ContractRemote

What you'll do

  • Instrument model, tool, retrieval, and end-to-end agent traces.
  • Build reliable pipelines for quality, token, cost, and latency data.
  • Create scorecards that help teams find regressions and expensive failure patterns.
  • Define alerting and data-quality checks for production evaluation systems.

What you bring

  • Experience with distributed tracing, telemetry, or production data systems.
  • Strong SQL plus proficiency in Python or TypeScript.
  • Ability to turn noisy event data into metrics people can act on.
  • Practical understanding of privacy and sensitive-data handling.
Bonus

Experience with OpenTelemetry, ClickHouse, data visualization, or LLM workloads.

TracingDataDashboards
Apply for this role
OPENING / 04

Technical Solutions Lead

Lead technical audits and turn ambiguous AI reliability problems into clear, measurable delivery plans.

Full-timeRemote

What you'll do

  • Run technical discovery with engineering, product, and operations leaders.
  • Map agent workflows, failure modes, model spend, and business success criteria.
  • Shape evaluation and improvement programs with concrete milestones.
  • Keep technical work aligned with the outcome the client needs to measure.

What you bring

  • Experience leading complex software or AI delivery with external stakeholders.
  • Enough technical depth to discuss APIs, data flows, evaluations, and deployment trade-offs.
  • Excellent facilitation, scoping, and written communication skills.
  • A direct, evidence-led approach to decisions and expectations.
Bonus

Previous hands-on engineering experience or work in regulated environments.

DiscoveryAI strategyDelivery
Apply for this role

Hiring process

Direct, practical, transparent.

01

Introduction

A focused conversation about your work and what you want to own.

02

Working session

A practical discussion using a realistic problem—not a take-home project.

03

Team conversation

Meet the people you would work with and ask anything.

04

Decision

Clear feedback and next steps without a drawn-out process.

General application

Don't see your exact role?

Tell us what you would want to own and show us work that demonstrates it.

Email careers@cognosystech.com