How to Run an AI Timekeeping Pilot Your Lawyers Will Trust

An AI timekeeping pilot should prove a workflow, not confirm a sales claim. The question is not whether a product can find impressive examples during a demo. It is whether the firm can repeatedly turn real work into accurate, reviewable time entries without moving the burden from attorneys to the billing team.
That requires more than installing software for a few enthusiastic users. A credible pilot defines the decision in advance, uses a representative cohort, tests difficult work, measures errors as carefully as benefits, and gives the firm a clean way to stop.
Trust follows from that design. Lawyers can see what the system observes, what it misses, who can view the record, and where human judgment remains. Firm leaders get evidence they can inspect rather than a recovered-hours estimate they must simply accept.
1. Write the decision before choosing the cohort
Begin with a one-page pilot charter. It should name the operating problem, the capabilities being tested, the people responsible, and the decision the firm expects to make.
A useful objective is specific:
Determine whether the system can convert supported work into correctly attributed, useful, policy-compliant entries that timekeepers will review and billing staff will accept without additional rework.
Avoid objectives such as “test AI” or “increase innovation.” They provide no basis for deciding whether to proceed.
Assign four owners:
- An executive sponsor who can make the final go, extend, or stop decision.
- An IT or security owner for identity, integrations, data handling, and exit.
- A billing administrator who owns the entry-quality rubric and validates client rules.
- A participant representative who can raise usability and workplace-trust issues without going through the vendor.
Also record what would make the firm stop early: an unacceptable privacy failure, uncontrolled exports, repeated wrong-matter assignments, material integration failures, or review work above an agreed limit. A pilot with no stop condition is a staged rollout.
2. Choose a cooperative, representative group
The cohort needs people who will engage with onboarding and provide specific feedback. It should not consist only of technology enthusiasts or the firm’s best timekeepers.
Include a mix of roles, work styles, and timekeeping habits: a strong contemporaneous timekeeper, someone who often reconstructs a day later, and a constructive skeptic. For a litigation or insurance-defense firm, include work across different clients and matters so the pilot encounters different narrative preferences, codes, and outside counsel guidelines.
Cooperation matters because participants must review drafts, report mistakes, and attend short check-ins. Representativeness matters because a product that works only for the most enthusiastic users has not established a rollout case. Document why each participant is included, and report results by participant as well as in aggregate. An average can hide one unusable experience.
Hourglass is generally designed for firms with at least 50 attorneys, where timekeeping volume, billing administration, and guideline complexity make the operational change more consequential. A smaller firm with straightforward billing may reasonably conclude that a pilot’s implementation cost outweighs the available benefit.
3. Establish a fair baseline
Select a historical comparison period before anyone sees pilot results. Match it as closely as possible for matter mix, working days, leave, deadlines, and billing cycle. Record known distortions rather than quietly excluding them afterward.
The baseline might include:
- Time from work performed to entry approval.
- Approved hours and the share entered daily.
- Attorney and billing-team time spent reviewing or repairing entries.
- Matter, duration, narrative, and coding corrections.
- Known entry-level billing-rule defects.
- Participant sentiment about burden, transparency, and control.
Keep the stages of value separate. Observed activity is not necessarily billable work. A generated draft is not an approved entry. An approved entry is not necessarily billed, and a billed amount is not necessarily collected. A three-month pilot can test capture, draft quality, approval, and export far more directly than long-term collection outcomes.
The Hourglass ROI calculator can help a firm model directional impact, but the pilot scorecard should use the firm’s measured inputs and should not treat every observed minute as revenue.
4. Complete governance and security review before capture
Participants should not discover the data policy after installation. Before day one, explain in plain language:
- Which applications and activity sources are in scope.
- What content and metadata are observed and retained.
- What a timekeeper can pause, exclude, inspect, correct, or delete.
- What firm administrators, vendor personnel, and subprocessors can see.
- Whether customer data may be used to evaluate, personalize, or improve the vendor’s features or models.
- Where data is processed, how long each data type is retained, and what happens when the pilot ends.
Review applicable client restrictions, confidentiality duties, vendor-risk requirements, and professional rules. ABA Formal Opinion 512 discusses competence, confidentiality, communication, supervision, and fees in the use of generative AI. It is not a universal rule for every timekeeping system, so the firm should apply its own jurisdiction’s requirements and seek ethics advice where appropriate.
The voluntary NIST AI Risk Management Framework offers a useful structure: govern who is accountable, map the workflow and risks, measure performance and failures, and manage the rollout or exit.
For Hourglass, firms can review the current architecture and assurance summary on the security page. Product-specific diligence should still be reconciled with the firm’s contract and security materials.
5. Configure and test before asking lawyers to rely on it
Connect only approved sources. Validate identity and directory provisioning, client and matter mappings, billing-system fields, task and activity codes, and representative billing rules before measuring normal work. The integrations catalog identifies the applications and billing systems currently in scope for Hourglass.
Use a small test set with known answers. Include an easy matter, a person or domain associated with several matters, fragmented work, overlapping calendar and meeting evidence, a prohibited narrative term, and an intentionally incomplete record. Confirm that uncertainty can remain unresolved rather than being filled with a confident guess.
The billing administrator should verify any rules extracted from outside counsel guidelines before activation. If pilot feedback suggests a rule change, a billing administrator should approve it; one unusual correction should not silently become firm policy.
Test human control explicitly. A participant should be able to review and edit the proposed matter, duration, narrative, and codes. With Hourglass, a human must explicitly approve an entry before it is exported to the billing system. The firm should treat that approval step as a control to evaluate, not as friction to conceal.
6. Run normal work and separate calibration
Do not script the whole pilot. Once configuration tests pass, participants should use the system on ordinary daily work. Difficult, representative work is more informative than a polished demo scenario.
Mark the first part of the pilot as calibration. During that period, correct matter relationships, narrative preferences, codes, and rules; resolve access and integration issues; and re-onboard anyone whose initial setup was incomplete. Report calibration results separately so early learning does not inflate the stable error rate or get erased from the record.
A typical Hourglass pilot lasts three months. Hourglass strongly recommends onsite support throughout the first month for onboarding, observation, re-onboarding, support, and rapid improvements. Onsite participation is not mandatory, but firms should consider it seriously: prompt observation can distinguish a product limitation from a setup problem before either becomes a habit.
Hourglass generally checks in with the administrative team weekly and with pilot users weekly or every other week. Keep those meetings short and use a consistent issue log. Do not coach participants toward positive survey answers.
Be candid about coverage. Hourglass does not currently capture ordinary calls placed directly through iPhone or Android phone logs unless the activity passes through a supported provider, and it does not have a native Android capture application. Offline work, in-person conversations, paper review, and other unsupported activity may require a timer or manual entry. Measure those gaps rather than hiding them.
7. Use a scorecard the firm can audit
Agree on definitions, thresholds, and data owners before the measurement window. Medians and per-participant distributions are usually more useful than one cohort average.
| Dimension | Measure | Guardrail |
|---|---|---|
| Capture | Approved incremental activity against the documented baseline | Do not count every draft as billable |
| Timeliness | Median time from work to approved entry; percentage approved daily | Does not prove faster payment |
| Matter accuracy | Percentage retaining the proposed client and matter | Define split and multi-matter work |
| Narrative quality | Percentage approved without substantive edit; billing-team rubric | Style alone is not compliance |
| Duration quality | Percentage approved without adjustment; size of adjustments | Apply the firm’s increment policy |
| Compliance | Correct detections, false alerts, missed test violations, reviewer repairs | Requires verified rules and test cases |
| Review burden | Median active review minutes per user per day | Separate measured and self-reported time |
| Adoption | Active days and completed review or approval | Installation and login are not adoption |
| Trust | Pre/post survey of control, transparency, and willingness to continue | Firm should own the survey |
| Integration | Failed or rejected records, duplicates, latency, manual corrections | Test the firm’s actual configuration |
| Support | Response and resolution time; unresolved issues at close | Include setup and service quality |
Do not publish a vendor’s overall conversion or adoption figure without its cohort, period, denominator, definitions, and exclusions. Your own pilot result is more useful when the measurement method is visible.
8. Make one of four decisions
At the end, have the firm score the pre-agreed rubric before reviewing the vendor’s financial narrative:
- Proceed: critical thresholds are met, no stop condition occurred, and participants and billing staff support a phased rollout.
- Remediate and retest: a defined configuration or training issue has a credible fix and a bounded retest.
- Extend: the measurement period was not representative (for example, because of leave or an unusual matter mix), and the extension has one stated purpose.
- Stop: control, quality, coverage, adoption, or economics are not sufficient.
If the firm proceeds, expand in cohorts and continue monitoring review burden, entry quality, and adoption. Invoice deductions, realization, and collections require later billing-cycle data and should not be implied by pilot drafts alone.
If the firm stops, revoke access, disconnect integrations, identify what will be exported, and verify contractual retention and deletion duties in writing. The ability to leave cleanly is part of the product test.
A trustworthy pilot does not guarantee a purchase. It gives the firm enough evidence to make a reversible, defensible decision and gives lawyers a reason to believe that the system will support their judgment rather than replace it.