Skip to content

Day100Bench

A test case for AI in the first months after close.

Day100Bench is a data room from an invented add-on acquisition, an answer key, and a grading script. It asks the questions an integration lead gets in week one: what will hurt us, what we are paying for twice, which dates cannot slip, and what order the work should happen in. Run it on any model or vendor tool. We have not published model results yet.

Sample grading of one finding, the Pacific Timber consent issue. Citing section 14.2 with the right deadline counts. Citing page 4 counts. A wrong deadline gets half credit. Citing section 14.1 gets no credit. Citing section 19.4, which does not exist, costs two points. Below it, a timeline of the dates that cannot slip.

How an answer is graded

A finding counts only if it cites the right place.

This is the first issue in the answer key and the clause behind it. Below it are six ways a model might report it, and what the grader does with each. We ran each one through the grading script to confirm the result.

Pacific Timber Cooperative ASO Agreement, section 14.2

Pacific Timber is the target's largest client, at 14.2% of revenue. Its contract requires written consent to a change of control. If consent is not given within 30 days of close, the client can leave on 90 days' notice without penalty.

Change of Control. Any Change of Control of Administrator shall require the prior written consent of Client. If Administrator undergoes a Change of Control without such consent, Administrator shall notify Client in writing within five (5) business days, and if Client's written consent is not obtained within thirty (30) days following the Change of Control, Client may terminate this Agreement upon ninety (90) days' written notice without penalty and without regard to §3.1. "Change of Control" means any transaction or series of transactions by which a person or group other than Northwind Insurance Group, Inc. acquires control of, or all or substantially all of the assets of, the business unit performing the Services.
  • Change of control, citing section 14.2, deadline December 30, 2026, client can terminateCountsThis is the clause that decides it, with the right deadline and consequence. The clause sets two dates, notice by December 7, 2026 and consent by December 30, 2026, and either one is accepted. The issue is material, so it is worth 1.6 of the 25 points for finding issues.
  • Change of control, citing page 4 of the contractCountsThe page that holds the clause is also accepted.
  • Change of control, citing section 14.1No creditThe clause exists, but it covers assignment. It does not decide the change-of-control question.
  • Change of control, citing section 19.4Minus 2 pointsThe contract has no section 19.4. Every citation to a place that does not exist costs two points.
  • Right clause, but the deadline is 60 days after closeHalf creditThe clause gives 5 business days for notice and 30 days for consent. The finding earns half credit, and the run is flagged for a critical miss.
  • Right clause, deadline and consequence, but a wrong summaryCountsThe grader does not read free text or the recommended action. This is a known limit, explained below.

What the test asks

The questions an integration lead gets in week one.

The model gets all of the documents in one prompt, plus the list of candidate actions, and no other help. Each section below is graded by rule against the answer key, and each column in the score table matches one section.

What is going to hurt us?

35 points

The documents contain 22 issues. 9 are material, such as a change-of-control clause on the largest client and a stop-loss database kept on one analyst's desktop. The model has to name the type of each issue, cite where it sits, and for material issues give the deadline and what happens if nobody acts. A finding counts only when both the type and the citation match the answer key. A material finding with the wrong deadline or consequence gets half credit. One client contract looks like a consent problem but is not, and flagging it costs points.

What are we paying for twice?

15 points

Both companies' software schedules are in the room. The model has to find the products that appear on both, give the target's annual cost for each one, and add them up. Some entries look like matches but are not, and counting them costs points.

Which dates cannot slip?

10 points

There are 8 dates to track, from the client consent window to the day the seller stops hosting the claims system. Each needs the right date and a citation. One of them depends on an event the documents do not date, and the model has to say that instead of guessing.

What happens in what order?

25 points

The model gets 30 candidate actions. 14 are required and 3 are wrong for this deal. It has to choose, schedule them by day, respect notice periods and minimum gaps, meet the deadlines, and mark which actions cannot be undone. For example, a termination notice cannot be undone, but building a data map can.

The seller sends a new TSA. What changed?

15 points

In a second prompt the model gets an amended transition services agreement and an issue list. We run it two ways. In one, the list is the correct one from the answer key. In the other, the list is the model's own first answer, so an issue it missed stays missed. For each of 15 issues it has to say whether the issue is resolved, changed, worse, or unchanged, and cite the clause that decides it. The new draft also creates problems of its own.

What fails a run outright?

Reported separately

A run is flagged if it misses any of these: the Pacific Timber consent issue, with its deadline and consequence, the client consent date, the ClaimsLink non-renewal notice date, the claims platform cutover date. The flag does not change the score. It is shown in its own column next to the total.

The deal

A health plan administrator carved out of an insurance group.

Meridian Benefits Administrators, a PE-backed third-party administrator, buys Cascade Health Plan Services, a business unit carved out of an insurance group. The seller keeps running the claims platform, email, payroll, and EDI under a transition services agreement. The deal closes on November 30, 2026. Every company, person, contract, and number is invented, so nothing in the room is confidential.

13

documents in the data room

22

issues in them, 9 of them material

$209,284

a year of target spend on 7 products both companies already pay for

$172,032

of that spend cannot be cancelled on its own because of terms in the data room

8

dates that cannot slip

30

candidate actions, 14 required, 3 wrong for this deal

Overlapping spend is a list to investigate, not a savings figure

The two companies both pay for 7 products, and the target spends $209,284 a year on them. Terms in the data room stop $172,032 of that from being cancelled on its own.

  • Microsoft 365, $76,032. It is billed inside the transition services agreement, which cannot be ended one service at a time.
  • Verity Eligibility, $96,000. Moving it needs Verity's consent, and Verity can raise the price to list rates for the buyer.

Nothing in the documents blocks the other $37,252. The pack has no notice periods, migration costs, or data on whether the buyer can absorb the target's users, so that number is not a savings figure either. The test scores the overlap today. It does not yet score which spend can be removed, or when.

What is in the data room

DocumentFormatGiven in
Transition Services Agreement v1 (execution draft)documentFirst prompt
Amended and Restated Transition Services Agreement v2documentSecond prompt only
Pacific Timber Cooperative ASO AgreementdocumentFirst prompt
Bexley County Schools Administrative Services AgreementdocumentFirst prompt
RxBridge PBM Services AgreementdocumentFirst prompt
ClaimsLink EDI Clearinghouse AgreementdocumentFirst prompt
Verity Eligibility SaaS Subscription AgreementdocumentFirst prompt
Seller responses to diligence request listdocumentFirst prompt
Cascade system inventorytableFirst prompt
Meridian (buyer) software licence scheduletableFirst prompt
Cascade (target) software licence scheduletableFirst prompt
Cascade operating workbookworkbookFirst prompt
NW-Claims group export (47 groups)tableFirst prompt

Limits

What a score does not tell you.

Read these before you quote a number from this page.

  • The grader does not read free text or the recommended action. It checks the issue type, the citation, and for material issues the deadline and the consequence. An answer can pass those checks and still recommend the wrong action. We accepted this so that no model has to judge another one. Expert review of the recommended actions is still needed, and it is not part of the score yet.
  • This is one deal. It is a carve-out of a third-party administrator with a transition services agreement. A good score here says nothing yet about a software add-on or a manufacturing roll-up.
  • The documents are invented and clean. Real data rooms are larger and include scanned PDFs, missing pages, and conflicting versions. This one fits in a single prompt.
  • The plan section gives the model a list of 30 actions we wrote. It tests choosing, ordering, and timing, not whether a model can discover the work on its own. It also has no owners, headcount, approvals, or migration costs, because a data room does not contain them.
  • The name says 100 days, but the plan runs to the end of the transition services agreement, about nine months after close. Four of the eight dates fall in the first 31 days.
  • Three runs per model is a small sample. A gap of a few points between two models may not hold on a rerun.

Other benchmarks

How this differs from legal and finance benchmarks.

Descriptions are taken from each benchmark's own public write-up.

  • Harvey's Legal Agent BenchmarkMore than 1,250 legal tasks across practice areas, including M&A work such as change-of-control review. Graded against expert-written rubrics, and a task passes only if every criterion passes.
  • Vals AI CorpFinQuestions about long credit agreements. The test set is private.
  • LegalBenchAn open collection of short legal reasoning tasks.
  • Day100BenchOne post-close deal, graded by a script against a fixed key. It scores the order of the plan and the update after a contract is renegotiated. We did not find a public benchmark that scores either one. If you know of one, tell us and we will list it here.

Who built this

MigrateForce built Day100Bench.

We sell software for post-close integration work, so we have a stake in how AI tools look on this test. These are the safeguards.

  • The answer key was written before any model ran. Every change to the key, the documents, or the grading code gets a new version number, and scores from different versions are never mixed.
  • The grader is a Python script. No model is involved in grading.
  • Every raw model response is kept next to its graded report.
  • No one outside MigrateForce has reviewed the answer key yet. If you think an item is wrong, use the dispute link below.
  • Our own product goes on this page only after it runs the full test case. We will not estimate its score from other models' results.

Get the test case

The data room, the answer key, the prompts, and the grading code. Run any model or vendor tool on it and you will get the same number we would.

Request the test case

Grade your current tool

Send us what your tool produces on this data room. We grade it with the same script and send you the full report.

Send an output

Dispute an answer

If you have run an integration like this and think an item in the key is wrong, tell us which one and why. Accepted changes get a new key version and every model is rerun.

Dispute an item