AI Timekeeping for Law Firms: What It Captures, What It Misses, and How to Evaluate It

TimekeepingThe Hourglass Team
An open paper ledger with a deep-teal ribbon tab.

“Automatic timekeeping” can describe very different products. One may create an activity timeline for the lawyer to interpret. Another may suggest a duration but leave the client and matter blank. A third may prepare a complete draft entry with a narrative and billing codes.

That makes Is it automatic? the wrong first question. A better one is: Can the system turn the work it can observe into a reviewable billing record without pretending to know what it cannot?

To answer it, a firm needs to examine every step between an attorney doing the work and an approved entry reaching billing.

What AI timekeeping actually does

AI timekeeping usually starts with evidence already created during a digital workday: an email, calendar event, document edit, attended meeting, browser session, or supported phone call. It organizes that evidence and proposes one or more time entries.

There are eight distinct checkpoints:

  1. Work source: Where did the activity happen?
  2. Capture: Was the relevant activity observed?
  3. Duration: How much nonduplicative working time did it represent?
  4. Attribution: Which client and matter did it concern?
  5. Description: What narrative and billing codes does the evidence support?
  6. Compliance: Does the proposed entry follow the applicable client rules?
  7. Review: Did an authorized person inspect and approve it?
  8. Delivery: Did the approved entry reach the firm’s billing system?

A product can perform well at one checkpoint and poorly at another. Capturing a Word session does not prove that the correct matter was selected. Selecting the correct matter does not make the narrative accurate. An accurate narrative may still violate a client’s outside counsel guidelines.

Evaluate the chain, not the word “automatic.”

Three common approaches to capture

The boundaries vary by vendor, but most products rely on one or more of these approaches:

Application and system integrations

An integration can provide structured information from email, calendars, document systems, meeting platforms, phone providers, and other applications. Structured data can make an event easier to recognize, but only for sources the product supports and the firm authorizes.

Device or window activity

Software on a work device can record which supported applications, windows, documents, or sites are active. This may cover work that does not produce a calendar event or server-side record. The firm should understand whether the product observes application names, metadata, content, screenshots, or some combination, and which exclusions users and administrators control.

Content-aware processing

Some systems process the substance and context of work, not merely the name of an application. That can help distinguish two matters handled in the same program or describe how a document changed. It also creates more important confidentiality, access, retention, and model-processing questions.

No capture architecture is universally best. Broader observation can provide more context; narrower collection can reduce the amount of sensitive material processed. The right balance depends on the firm’s work, client commitments, security requirements, and tolerance for manual completion.

What a system can capture

A serious evaluation should begin with a maintained source inventory, not a claim that the software captures “everything.” Test the actual applications your lawyers use and the activities that matter most.

For example, Hourglass’s current application catalog includes Microsoft Outlook and Calendar, Word, Excel, PowerPoint, Teams and Teams Phone, OneNote, OneDrive, SharePoint, and Photos; Adobe Acrobat, Kofax, Foxit, and VLC; Zoom and Zoom Phone, RingCentral, and 8x8; Chrome, Edge, Brave, and Firefox; ChatGPT, Claude, Gemini, Harvey AI, Westlaw, and CoCounsel. The same page lists the billing systems currently supported for delivery of approved entries.

Hourglass uses rich activity evidence to construct a chronological account of work. For example, it may show that a timekeeper opened and reviewed part of a document, made changes, joined a meeting, or worked in another supported application. That evidence can inform duration, matter attribution, narrative, and code suggestions. The Hourglass timekeeping page explains the product workflow at a higher level.

This is vendor-specific behavior, not a definition of the whole category. Every firm should create its own test matrix using its actual software, devices, permissions, and working patterns.

What AI timekeeping can miss

Digital capture is not a complete account of legal work. Common gaps include:

  • In-person conversations, handwritten analysis, and work performed away from an observed device.
  • Activity in unsupported, excluded, paused, private, or offline applications.
  • Ordinary mobile calls when no supported phone provider supplies the event.
  • A new matter whose people, documents, or systems have not yet established enough context.
  • Work whose evidence could plausibly belong to more than one matter.
  • The lawyer’s purpose, judgment, or result when the source record does not support an inference.

Hourglass, for example, does not capture ordinary calls placed directly through an iPhone or Android handset unless the activity passes through a supported service such as Teams Phone, Zoom Phone, RingCentral, or 8x8. Users can add missing work manually, use a smart timer that tracks the work performed during the timed session, or delegate entry review to an authorized assistant.

The important behavior is not whether a product ever encounters ambiguity. It is what the product does next. When Hourglass lacks enough evidence for a matter, narrative, task code, or activity code, it can leave that field blank for the user instead of manufacturing certainty.

“Accuracy” is six different questions

A single accuracy percentage hides the mistakes a firm actually needs to find. Measure at least six dimensions:

  1. Coverage: What percentage of known work events did the system observe?
  2. Duration: Were working minutes reconciled without idle time, overlap, or duplication?
  3. Attribution: How often were the client and matter correct, incorrect, or explicitly uncertain?
  4. Description: Were narratives accepted, lightly edited, substantially rewritten, or rejected?
  5. Compliance: Did the product identify applicable billing issues without overwhelming users with false positives?
  6. Release control: Did every exported entry receive the required human approval?

Duration deserves particular scrutiny. A calendar event and a Zoom attendance record may describe the same meeting, not two billable activities. Hourglass correlates evidence for the same event to avoid double-counting, but the resulting duration remains subject to timekeeper review.

Matter attribution should also expose uncertainty. Hourglass can associate people, email domains, organizations, locations, and previously approved work with a client or matter. Those relationships support a suggestion; they do not eliminate the need for review.

Human review is a control, not a failure

An AI-generated billing entry is a draft. The timekeeper should be able to inspect the underlying record, correct the duration and matter, edit the narrative and codes, split or merge entries, reject a suggestion, and add missing work.

In Hourglass, drafts cannot reach the billing system without explicit human approval. The audit record retains original evidence, the model suggestion, user edits, approvals, overrides, and export history. Approved entries then flow to the firm’s supported billing system; Hourglass does not claim to observe every downstream prebill change, deduction, appeal, or payment.

That boundary matters beyond product design. The ABA’s Formal Opinion 512 explains that lawyers using generative AI should understand a tool’s capabilities and limitations and apply an appropriate degree of independent review. It also addresses confidentiality and reasonable fees. The applicable duties depend on the jurisdiction, matter, client agreement, and use case; an AI product does not discharge them for the lawyer.

Evaluate privacy as a data flow

A certification is useful evidence, but it does not answer every confidentiality question. Map the path of information:

  1. What source activity or content is collected?
  2. Where is it processed and stored?
  3. Which firm users, vendor personnel, and subprocessors can access each layer?
  4. Is customer data used to personalize, evaluate, or improve the product?
  5. What can users pause, exclude, inspect, correct, or delete?
  6. What retention terms apply to core records, model inputs and outputs, diagnostics, logs, and backups?
  7. What happens to the data when the agreement ends?

The Hourglass security overview describes its US-hosted, firm-specific application deployment and database, encryption, SOC 2 Type 2 examination, access controls, and processing-data retention. Buyers should still reconcile public material with their security review and operative agreement.

This inquiry follows the same underlying concern as ABA Model Rule 1.6: lawyers must make reasonable efforts to prevent unauthorized disclosure of or access to information relating to a representation. This article is an evaluation framework, not legal or ethics advice.

Run a pilot against the firm’s real work

A polished demonstration cannot reveal how a product handles your matters, client rules, and exceptions. A useful pilot should include a representative mix of timekeepers, not only enthusiasts, and ordinary litigation or transactional work.

Agree on the scorecard before the pilot begins:

  • Percentage of known work events observed.
  • Percentage of minutes reconciled without overlap or duplication.
  • Correct, incorrect, and uncertain matter-attribution rates.
  • Narratives accepted unchanged, lightly edited, or rewritten.
  • Valid and false-positive billing-compliance flags.
  • Median review time and time from work to approved entry.
  • Unsupported or off-screen work added manually.
  • Successful delivery to the billing system and handling of failures.
  • Adoption and active use across different timekeeper profiles.
  • Security and data-flow questions resolved before production use.

Do not treat every observed minute as billable or every suggestion as recovered revenue. A defensible business case begins with entries that timekeepers actually approve and the firm can appropriately bill.

The final evaluation checklist

Before choosing a system, ask the vendor to demonstrate:

  • The exact capture inventory and known gaps.
  • Duplicate, idle-time, interruption, and uncertain-attribution handling.
  • The evidence behind narratives and billing codes.
  • A safe refusal or blank-field behavior when evidence is insufficient.
  • Client-specific compliance using the firm’s real guidelines.
  • Complete timekeeper control and explicit approval before export.
  • The billing-system connection, supported fields, and failure behavior.
  • The data flow, access model, retention terms, exclusions, and deletion process.
  • Pilot results separated by coverage, duration, attribution, narrative, compliance, and adoption.
  • The methodology and customer permission behind every performance claim.

The best AI timekeeping system is not the one that promises to know the entire workday. It is the one that captures enough of the right evidence, turns it into useful drafts, clearly identifies uncertainty, and leaves the final billing decision with the people responsible for it.

All posts