Benchmark
Add-On Integration Readiness Benchmark.
One job, scored. A platform company has just closed an add-on. The model gets the data room and has to say what will bite, where the document says so, what the licence overlap is worth, which dates cannot slip, and in what order the first hundred days run. Then the seller changes one agreement and the model has to say what moved.
Scores
No scores are published yet.
The benchmark pack is not present in this build.
Method
Five sections, one hundred points, no judge model.
Matching is by category and citation. Every citation is checked against the pack, so a made-up clause number is detected without a human in the loop. The plan and changed-source sections are the two that separate this from a document question-answering test.
Findings, 35 points
Every material issue in the data room, each cited to a clause, page, row, or cell. A finding scores only if the category is right and the citation is in the answer key. Text is never compared.
Licence overlap and gates, 25 points
The same-product overlap between the buyer and target licence schedules, its target-side dollar total, and the dated deadlines the integration has to hit, computed from the close date and the documents.
Plan, 25 points
An ordered first-hundred-days plan from a step catalog that includes steps that are wrong for the deal. Scored on required steps, precedence, minimum gaps, deadlines, and whether each step is correctly labelled reversible or not.
Changed source, 15 points
The seller renegotiates the transition services agreement. Given the accepted findings and the new document, mark each finding unchanged, changed, resolved, or worse, citing the clause that decides it, and list what is new.
Get the pack, the key, and the grader
The full pack tpa-001, its answer key, the two prompts, and the scorer, so you can run any model or any vendor against it and get the same number we would.
Request the packHave your current tool graded
Send us the output your AI tool or vendor produces on this pack and we grade it with the same scorer.
Send an output for gradingThis is the job MigrateForce is built to run. Our own score appears on this page when our pipeline has run the pack end to end, not before.