How to Evaluate AI Timekeeping Software for a Litigation Firm

TimekeepingThe Hourglass Team
A magnifying glass resting over a blank paper document.

A litigator’s day is a poor fit for a neat activity log. Drafting is interrupted by email. One adjuster appears across several matters. A short call changes the direction of a motion. A deposition or hallway conversation leaves little desktop evidence at all. Then every resulting entry must satisfy a client’s narrative rules, codes, and outside counsel guidelines.

That is why a feature checklist is not enough to evaluate AI timekeeping software for a litigation firm. The real test is whether the system can turn a fragmented day into accurate, defensible entries that require little correction and arrive safely in the firm’s billing system.

The best way to find out is a scored pilot built around the firm’s actual work.

Define success as an approved entry, not detected activity

An application can detect a document opening or meeting attendance without producing a usable time entry. A complete evaluation should examine how the product observes work, suggests a matter, drafts the entry, checks the billing rules, obtains timekeeper approval, and delivers the approved record to the billing system.

Each transition can fail independently. A tool might observe most desktop activity but assign it to the wrong matter. It might identify the right matter but produce a vague narrative. It might draft a strong narrative but apply the wrong client code. Measuring only “time captured” hides those differences.

Before seeing a demo, define the unit of success as an entry that is accurate, appropriately coded, compliant with the applicable rules, approved by a person, and delivered to the billing system. Then evaluate the steps that create it.

Seven dimensions to evaluate

1. Capture coverage and declared blind spots

Ask each vendor for an exact inventory of supported applications and capture methods. “Automatic” might mean observing desktop activity, importing calendar events, interpreting documents, or simply suggesting narratives from timer notes.

Test the tools where litigators actually work: email, word processing, PDFs, research sites, videoconferences, calendars, supported phone systems, and the firm’s billing platform. Hourglass publishes its current scope in its application and billing-system catalog.

Just as important, test what happens away from a supported source. Ordinary iPhone or Android calls that do not pass through a supported business phone provider are a known Hourglass blind spot, and Hourglass does not currently offer native Android capture. Depositions, court appearances, travel, paper review, and in-person conversations may also lack enough digital evidence. The evaluation question is therefore not whether a vendor claims to capture everything. It is whether missing work is visible and easy for a timekeeper to add correctly.

2. Matter attribution under real ambiguity

Litigation matter attribution is harder than matching a document name. The same client, adjuster, expert, opposing counsel, or email domain can appear across several active cases. A credible pilot should deliberately include those collisions.

Hourglass uses relationships among people, domains, organizations, locations, and prior approved work to suggest a client and matter. If the evidence is insufficient, it can leave the field blank instead of making a confident-looking guess. Test both behaviors: the rate of correct first suggestions and whether uncertainty is exposed at the right time.

Do not let a vendor choose only clean examples. Include a shared document, a thread spanning two matters, and work whose purpose becomes clear only after another event. Record how many attributions the timekeeper must change and whether corrections improve later suggestions without crossing client or matter boundaries.

3. Duration and grouping quality

Raw application time is not billable time. A calendar event and Zoom attendance may describe one meeting, not two activities. A document may remain open while the lawyer handles another matter. Several short sessions might properly become one entry or need to remain separate to avoid block billing.

Ask vendors to explain their publishable method for idle time, interruptions, overlap, rounding, and double-counting. Then compare draft duration against a corroborated record for a small set of authorized test days.

Hourglass correlates evidence that describes the same event and can group related work, such as reviewing a document before writing an email about it. The result is still a draft. The timekeeper can edit, split, merge, reject, or add to it before approval.

4. Narratives, codes, and client rules

For insurance-defense firms, capture accuracy without compliance accuracy is an incomplete result. Score whether the draft describes a discrete task, contains only supported facts, selects an allowed task and activity code, and survives the client’s rules without substantive rewriting.

“Supports UTBMS” is not specific enough. The LEDES Oversight Committee’s UTBMS resource explains that task codes organize work by phase while activity codes describe the type of service. Clients can require different code sets and levels of detail. Test the firm’s actual allowed combinations, not a generic code list.

In Hourglass, available codes are defined by the firm and synchronized from the billing system. Uploaded guidelines become proposed rules that a person must verify and approve. Supported entry-level rules can warn, explain, suggest a correction, or block release according to administrator configuration. The current scope does not include expense, rate, or budget controls, nor controls that require analysis across an invoice or a matter’s full history.

5. Review burden and human control

The right comparison is not generated output versus a blank page. It is the total work needed to reach an approved entry.

Have pilot participants review drafts during ordinary days and record the time spent correcting matter, duration, narrative, and codes. Separate cosmetic edits from substantive ones. Also test rejection, manual additions, delegation, and a deliberately weak-evidence example.

Every Hourglass entry remains under human control. A person must explicitly approve and export it before it reaches the billing system. The audit record retains original evidence, the model suggestion, edits, approvals, overrides, and export history. Firms seeking autonomous, unreviewed billing should not treat Hourglass as that product.

6. Integration depth and failure handling

“Integrates with” can mean anything from a file export to a maintained exchange of matters, timekeepers, codes, and entries. Require a demonstration using the firm’s actual billing environment.

For Hourglass, the verified boundary is clear: billing-system configuration data needed for the firm environment flows into Hourglass, and approved time entries flow back out. Hourglass does not claim that later invoice deductions, appeals, collections, or write-offs automatically return as structured data.

During the pilot, test a normal submission, an edited entry, a temporary connection failure, and a retry. Ask who sees an error, how duplicates are prevented, and what evidence confirms receipt.

7. Security, privacy, and governance

Timekeeping evidence can contain confidential information relating to a representation. The commentary to ABA Model Rule 1.6 makes clear that confidentiality extends beyond privileged communications. ABA Formal Opinion 512 also discusses competence, confidentiality, supervision, communication, and reasonable fees in lawyers’ use of generative AI. A firm should involve its own security, privacy, and professional-responsibility owners in the evaluation.

Ask what content is collected, where it is processed, who may access it, how long each data layer is retained, which third parties receive it, and whether it is used to improve the vendor’s models. Request contractual answers, not only marketing language.

Hourglass is US-hosted, has completed a SOC 2 Type II examination, and gives each firm its own application deployment and database. Restricted third-party model endpoints operate under zero-data-retention terms and do not use customer data to train their own general-purpose models. Hourglass may use customer data, generated entries, corrections, approvals, and feedback to improve Hourglass features and models. More detail is available on the Hourglass security page and through customer diligence.

Build a litigation-shaped pilot

Use appropriately authorized matters and include ordinary work, not only staged demonstrations. A useful test set contains:

  • two active matters with overlapping people or organizations;
  • drafting interrupted by email, research, a call, and another matter;
  • several short sessions that should become one defensible entry;
  • related work that should stay separate under a block-billing rule;
  • a hearing, deposition, travel period, or in-person conversation with weak digital evidence;
  • different client guideline and coding requirements; and
  • a failed billing-system submission followed by recovery.

Capture a baseline before the pilot. For a representative sample, document how entries are created today, how long review takes, how often matters and codes change, and which guideline problems billing staff find before submission. Use the same definitions during the pilot.

Then score each product consistently:

Dimension Example measure
Capture completeness Share of known billable activities represented
Matter attribution Share correct before user correction
Duration Variance from the corroborated test-day record
Narrative quality Share accepted without substantive rewriting
Coding and rules Entry-level issues found before approval
Review burden Median review and edit time per user per day
Integration Successful submissions and recoverable failures
Adoption Active use and approved entries within the cohort
Trust User and administrator confidence, recorded consistently

Set the firm’s own pass, remediation, and failure thresholds before the pilot starts. There is no defensible universal accuracy or adoption percentage for every litigation practice. If downstream write-down or collection data matters to the business case, measure it separately from the billing or finance system and allow enough time for the relevant billing cycle. Do not assume the timekeeping product observes those outcomes.

Ask for evidence, not adjectives

Before choosing a vendor, ask for:

  1. A current capture inventory and explicit list of unsupported work.
  2. A demonstration of ambiguity, abstention, and user correction.
  3. The firm’s actual code set and client rules tested against sample entries.
  4. A live integration workflow, including a failed submission and recovery.
  5. A data-flow explanation, retention schedule, subprocessor terms, and available assurance reports.
  6. Pilot support commitments and named owners on both sides.
  7. The methodology behind every accuracy, adoption, revenue, or time-saved claim.

For a directional business case, the Hourglass impact calculator lets a firm edit its own timekeeper counts, rates, annual hour requirements, and e-billing deductions. It is not a substitute for a measured pilot.

The practical fit test

Hourglass is built primarily for mid-sized and larger firms, generally those with 50 or more attorneys, where billing volume, fragmented work, and client-rule complexity make capture and compliance operational problems. It is particularly well suited to insurance-defense workflows. A smaller firm with simple billing may not realize enough value to justify the implementation and organizational change.

The decisive question is not which product detects the most activity. It is which one produces the most accurate, defensible, compliant entries with the least total correction while keeping the lawyer accountable for what is ultimately billed. A pilot designed around the difficult parts of litigation will answer that question far more reliably than any vendor comparison table.

All posts