Skip to content

Benchmark

Add-On Integration Readiness Benchmark.

One job, scored. A platform company has just closed an add-on. The model gets the data room and has to say what will bite, where the document says so, what the licence overlap is worth, which dates cannot slip, and in what order the first hundred days run. Then the seller changes one agreement and the model has to say what moved.

Scores

No scores are published yet.

The benchmark pack is not present in this build.

Method

Five sections, one hundred points, no judge model.

Matching is by category and citation. Every citation is checked against the pack, so a made-up clause number is detected without a human in the loop. The plan and changed-source sections are the two that separate this from a document question-answering test.

Findings, 35 points

Every material issue in the data room, each cited to a clause, page, row, or cell. A finding scores only if the category is right and the citation is in the answer key. Text is never compared.

Licence overlap and gates, 25 points

The same-product overlap between the buyer and target licence schedules, its target-side dollar total, and the dated deadlines the integration has to hit, computed from the close date and the documents.

Plan, 25 points

An ordered first-hundred-days plan from a step catalog that includes steps that are wrong for the deal. Scored on required steps, precedence, minimum gaps, deadlines, and whether each step is correctly labelled reversible or not.

Changed source, 15 points

The seller renegotiates the transition services agreement. Given the accepted findings and the new document, mark each finding unchanged, changed, resolved, or worse, citing the clause that decides it, and list what is new.

Get the pack, the key, and the grader

The full pack tpa-001, its answer key, the two prompts, and the scorer, so you can run any model or any vendor against it and get the same number we would.

Request the pack

Have your current tool graded

Send us the output your AI tool or vendor produces on this pack and we grade it with the same scorer.

Send an output for grading

This is the job MigrateForce is built to run. Our own score appears on this page when our pipeline has run the pack end to end, not before.