Tech-Horizonz

Independent AI quality lab

AI agents, tested like production software.

Tech-Horizonz builds LLM agents for software quality work and evaluates them with the rigor used on regulated financial systems: repeatable scenarios, automated checks, and evidence you can trace.

R&D stage · prototypes in active development

eval · conversation rungpt-4.1-mini
scenario
refund-request-07
agents
customer ↔ support
turns
12 / 12
  • stays in rolepass
  • policy followedpass
  • no invented factspass
  • tone · llm judgereview
transcript + report saved3/4
Illustrative output. Shows how a run is reported.

Focus

Four problems we work on

Where AI can speed up software quality work, and where it most needs checking.

01 / evaluate

LLM conversations

Repeatable scenarios scored by rule-based checks and an LLM judge, with full transcripts.

02 / triage

Test failures

Test results exposed to an LLM through MCP, so failed runs come back grouped by likely cause.

03 / report

QA status

Daily engineering reports where systems supply the facts and the model only writes the summary.

04 / automate

Agent workflows

Agents that turn business rules into test assets, returning structured JSON for downstream automation.

Projects

What's on the bench

Current status is shown on each project. Code is private; walkthroughs are available on request.

Working prototype

LLM Conversation Tester

Two agents with distinct roles and system prompts run 20+ repeatable scenarios on GPT-4.1-mini. Every transcript is scored with automated checks and LLM-as-judge evaluation, and failures produce a report.

  1. Scenario
  2. Two agents
  3. Checks + judge
  4. Failure report

Python · OpenAI API · LLM-as-judge · JSON transcripts

In development

MCP Playwright Triage

An MCP server that gives an LLM access to Playwright test results for failure triage. Proposed explanations are treated as hypotheses tied to test evidence, not established root causes.

  1. Playwright run
  2. MCP server
  3. LLM agent
  4. Evidence-linked triage

TypeScript · Model Context Protocol · Playwright · JSON reporter

In design

AI QA Daily Report

Deterministic test and defect systems calculate the facts. The LLM summarizes only controlled, structured inputs, and its draft is checked for unsupported claims before anyone reads it.

  1. Test + defect data
  2. Calculated facts
  3. LLM summary
  4. Claim check

Python · Structured JSON output · Tool calling

Experimental

Test Automation Agents

Python and TypeScript agent workflows that connect business rules to backend services to help create and maintain tests, with structured JSON outputs that stay machine-readable.

  1. Business rules
  2. Agent + tools
  3. Validated JSON
  4. Test assets

Python · TypeScript · LangChain · Guardrails · LangSmith

Approach

How the work is held to account

  1. 01

    Evidence over eloquence

    A model's explanation is a hypothesis until a test result or log backs it up.

  2. 02

    Systems count, models write

    Numbers come from systems of record. The LLM summarizes; it doesn't calculate.

  3. 03

    Structured by default

    Schema-checked JSON outputs, guardrails, and traces on every agent.

  4. 04

    Every agent ships with evals

    Scenarios, rule-based checks, and LLM-as-judge scoring, rerun whenever the agent changes.

PythonTypeScriptPlaywrightMCPLangChainLangSmithOpenAI APIChromaNext.jsFirebase

Founder

Ahmed Elsabban

Senior SDET and test automation lead with 8+ years in regulated financial services and enterprise utilities. Tech-Horizonz applies that quality-engineering background to AI.

At FINRA he built a Cypress/TypeScript framework from scratch to 1,500+ automated tests and led a six-person QA team.

founded
2025
based in
Northern Virginia, USA
background
FINRA · S&P Global · Duke Energy
stage
Research and development

Contact

Evaluating AI agents? Let's compare notes.

Questions about the projects, AI quality engineering, or a walkthrough of the work.

info@tech-horizonz.com