← Sottava, jobs the hour they open
2 mo agofound 2 h ago
Head of Evaluations (Legal AI Benchmarking)
Posted by Newcode.ai on 15 July 2026, 87 days ago. Still on their Workable board when we checked 1 min ago.
Read out of the posting
LevelNot stated
Experience askedNot stated
EmploymentNot stated
LocationUnited States
RemoteYes
Visa sponsorshipNot stated
SalaryNot published, and most postings do not
Posted2026-07-15
Found viaworkable, direct from their system
We saw it 3 months after it went up.
The posting, as the company wrote it
Who are we?
At Newcode.ai, we're transforming how law firms and legal professionals harness AI for real-world impact. As part of our collaborative, high-growth team, you'll have the rare opportunity to work side-by-side with visionary founders at the bleeding edge of AI and legal innovation — shaping not just our product, but the future of legal work itself.
Note: We believe in being transparent about what it's like to work at Newcode. As a fast-growing startup, we're building and evolving every day. That means not every process, playbook, or framework is already in place, and priorities can shift quickly.
The people who thrive here are comfortable with ambiguity, take ownership, and don't wait for perfect direction. They are resourceful, proactive, and able to "figure it out"—solving problems, creating structure where needed, and helping build the company as they go. If you are good with this then, great! Keep reading to learn more.
Position Overview
We are seeking a highly analytical professional with a strong statistical background to join our Head of Evaluations. In this role, you will design, implement, and scale the testing frameworks used to evaluate our platform. You will ensure our AI products meet the highest standards of legal reasoning, factual accuracy, and regulatory compliance while maintaining a near-zero hallucination rate.
Key Responsibilities
Design Legal Benchmarks for: Contract Drafting, Information Extraction, Legal Research, and Contract Review
Build, source and maintain relevant datasets
Audit AI Output: Review and score complex AI-generated legal text, contract analyses, and statutory interpretations for accuracy and precision and lay out a strategy.
Define Evaluation Metrics: Establish clear criteria for grading model performance, specifically focusing on logical reasoning, citation accuracy, and the model's ability to safely abstain from answering.
Collaborate with Engineering: Partner directly with Engineering to translate legal errors into actionable technical feedback for model fine-tuning.
Requirements
PhD or Masters in statistics, mathematics, machine learning or equivalent
Analytical Skills: Proven ability to break down complex statutory frameworks and case law into structured, logical data points.
Tech-Savviness: python, panda, numpy, jupiter notebooks and similar statistical models
Visa Sponsorship:
At this time, Newcode is unable to provide visa sponsorship. Candidates must be authorized to work in the applicable country without employer sponsorship.
Copied from Newcode.ai’s own board, not rewritten. Original ↗
Also open at Newcode.ai
Why this page exists
We read companies’ own hiring systems every hour, 1,769 of them, and show a job the hour it opens instead of when a job board gets around to indexing it. We saw it 3 months after it went up.
The feed is free. No card, no trial to expire.
Apply at Newcode.aiA free account first, no card