LLM conversations
Repeatable scenarios scored by rule-based checks and an LLM judge, with full transcripts.
Independent AI quality lab
Tech-Horizonz builds LLM agents for software quality work and evaluates them with the rigor used on regulated financial systems: repeatable scenarios, automated checks, and evidence you can trace.
R&D stage · prototypes in active development
Focus
Where AI can speed up software quality work, and where it most needs checking.
Repeatable scenarios scored by rule-based checks and an LLM judge, with full transcripts.
Test results exposed to an LLM through MCP, so failed runs come back grouped by likely cause.
Daily engineering reports where systems supply the facts and the model only writes the summary.
Agents that turn business rules into test assets, returning structured JSON for downstream automation.
Projects
Current status is shown on each project. Code is private; walkthroughs are available on request.
Two agents with distinct roles and system prompts run 20+ repeatable scenarios on GPT-4.1-mini. Every transcript is scored with automated checks and LLM-as-judge evaluation, and failures produce a report.
Python · OpenAI API · LLM-as-judge · JSON transcripts
An MCP server that gives an LLM access to Playwright test results for failure triage. Proposed explanations are treated as hypotheses tied to test evidence, not established root causes.
TypeScript · Model Context Protocol · Playwright · JSON reporter
Deterministic test and defect systems calculate the facts. The LLM summarizes only controlled, structured inputs, and its draft is checked for unsupported claims before anyone reads it.
Python · Structured JSON output · Tool calling
Python and TypeScript agent workflows that connect business rules to backend services to help create and maintain tests, with structured JSON outputs that stay machine-readable.
Python · TypeScript · LangChain · Guardrails · LangSmith
Approach
A model's explanation is a hypothesis until a test result or log backs it up.
Numbers come from systems of record. The LLM summarizes; it doesn't calculate.
Schema-checked JSON outputs, guardrails, and traces on every agent.
Scenarios, rule-based checks, and LLM-as-judge scoring, rerun whenever the agent changes.
Founder
Senior SDET and test automation lead with 8+ years in regulated financial services and enterprise utilities. Tech-Horizonz applies that quality-engineering background to AI.
At FINRA he built a Cypress/TypeScript framework from scratch to 1,500+ automated tests and led a six-person QA team.
Contact
Questions about the projects, AI quality engineering, or a walkthrough of the work.
info@tech-horizonz.com