Building a compliant product is just the first step. But your success as a vendor depends on delivering what your customers want: an effective product.
Vendor Build Support
AI has enabled assessment formats that were previously impractical: large-scale free-text scoring, chat-based interviews, high-volume role plays, adaptive exercises, rapid item generation, and personalised candidate experiences.
You need scoring that is robust and fair. You need to know what to trial, and what a successful trial looks like. You need features in place now to meet new obligations under the EU AI Act. And you need assessment experts and developers working together to achieve your product roadmap.
What we do
We work with assessment vendors before build, during build and after launch.
We focus on the assessment underneath the product: what the tool measures, how the score is produced, how the evidence is tested, and what claims you can justifiably make.
Construct definition
Rubric and LLM prompt design
Scoring architecture
Synthetic data generation and test-case design
Adversarial testing of scoring behaviour
LLM calibration and consistency checks
Trial and pilot design
Reliability and validity evidence planning
Fairness and adverse-impact analysis
Regression testing
Buyer-facing documentation and RFP response
Where we help
Design
1
We review what your product claims to measure and whether your data supports that claim.
That includes construct definition, role relevance, candidate experience, decision role and the claims your sales team can safely make.
Build
2
We review rubrics, prompts and scoring logic, working in partnership with both your assessment experts and developers.
This is especially important for LLM-scored assessments. The tech team - rightly - owns the implementation, but assessment experts still need to own the scoring criteria, calibration logic and interpretation rules.
We can help test whether the LLM follows the rubric, whether it overscores fluent answers, whether it assigns similar scores to similar evidence, and whether small prompt changes shift the score distribution.
Pilot and test
3
We design tests that deliberately try to break the scoring.
That can include synthetic candidate responses, edge cases, adversarial examples, prompt-variation tests and subgroup checks. The aim isn’t to prove the product is perfect. The aim is to find the failures before a buyer, regulator or a candidate does.
Typical LLM pitfalls include:
● Scoring fluency instead of the target behaviour
● Treating longer answers as better answers
● Rewarding generic leadership language
● Missing good evidence when candidates use unusual phrasing
● Giving plausible explanations for inconsistent scores
● Changing scoring behaviour after a prompt or model update
● Applying the rubric differently across demographic or language groups
● Producing confidence where the evidence is thin
Validate
4
We help design the evidence plan.
That may include pilot success criteria, reliability checks, validation study design, adverse impact analysis, local-use evidence, and defining the limitations of how the assessment should be used.
AI assessment products change. Prompts change. Models change. Data sources change. Scoring rules change. New customers use the product in new roles or countries.
We help define what needs monitoring, what needs re-testing, and what should be documented before the next renewal, expansion or procurement review.
Post-launch
5
Where we’ve done it.
We’ve applied the same discipline to products and markets including:
● Conversational and chat-based assessment.
● Early careers and volume hiring.
● Assessment for development.
● In-house tools built by employers for their own use.
How we work
We work quickly and focus on what matters. Whatever stage you’re at in your build, we’ll tell you what we need - usually your build documentation, a system demo, and time with selected technical and non-technical stakeholders.
You get a focused written report, usually within two weeks. It covers risks, remediations, design points, testing and trialling needs, and anything else relevant to your build. We always include free follow-up calls to review the report, answer questions, and give further advice.