I spent last week at a legal technology conference. I am not a lawyer, but e-discovery has become one of those areas where the security, IT, and legal worlds collide in ways that keep all three groups uncomfortable. After two days of presentations and vendor demos, my primary takeaway is this: e-discovery is genuinely, deeply, almost absurdly hard, and the gap between what the law requires and what most organizations can actually deliver is enormous.
The Scale Problem
The numbers are staggering. A single corporate custodian - one employee whose electronic communications are subject to discovery - typically has between 5 and 15 gigabytes of email alone. A mid-sized litigation matter might involve 50 to 100 custodians. That puts you at 250 GB to 1.5 TB of email before you even touch file shares, instant messages, voicemail, SharePoint, or the growing universe of SaaS applications where business data now lives.
The Sedona Conference estimates that the average cost to review one gigabyte of electronic data is between $18,000 and $30,000 when you account for attorney review time. Do the math on a matter involving a terabyte of potentially responsive data and you arrive at numbers that make CFOs physically ill.
And the data volumes are only growing. Enterprise email archives are larger this year than last year. Employees are using more applications, generating more data, across more platforms. The problem does not get easier with time.
Preservation Obligations
The legal duty to preserve potentially relevant evidence kicks in the moment litigation is reasonably anticipated - not when a lawsuit is filed, but when someone in the organization could reasonably foresee that a dispute might lead to litigation. That trigger point is often much earlier than most IT departments realize.
Once the duty attaches, the organization must issue a litigation hold - a directive to preserve relevant documents and suspend any routine deletion policies that might destroy evidence. This sounds straightforward until you try to implement it across a modern enterprise. How do you ensure that automated email archiving policies do not delete relevant messages? What about backup tape rotation? What about the data on the laptop of a traveling sales rep who has not connected to the corporate network in two weeks?
Failure to preserve carries serious consequences. Courts have imposed sanctions ranging from adverse inference instructions - essentially telling the jury to assume the destroyed evidence was unfavorable - to default judgments. The Qualcomm case in 2008, where the court found the company had failed to produce hundreds of thousands of relevant emails, resulted in sanctions against the company and its outside counsel. That case sent a chill through both the legal and IT communities.
Technology Solutions and Their Limits
The e-discovery technology market has exploded over the past five years. Early case assessment tools, predictive coding, concept clustering, de-duplication engines, email threading algorithms - there is no shortage of technology designed to reduce the volume of data that requires human review.
Predictive coding in particular has generated significant interest. The idea is to use machine learning to train a model on a subset of documents that have been manually reviewed, then apply that model to the full collection to identify likely relevant documents. Early results are promising - some studies suggest that predictive coding can achieve accuracy comparable to or better than exhaustive manual review at a fraction of the cost.
But the technology adoption curve in legal is notoriously slow. Many courts have not yet addressed the admissibility of predictive coding results. Many law firms are reluctant to rely on automated review for fear that opposing counsel will challenge the methodology. And the technology itself requires careful implementation - garbage in, garbage out applies to training sets just as much as it applies to any other machine learning application.
What IT and Security Can Do
If you are on the IT or security side of the house, here is what I would recommend. First, know where your data is. Maintain a current data map that identifies the types, locations, and retention characteristics of electronically stored information across the organization. Second, have a defensible retention policy and actually enforce it. Keeping everything forever is not a strategy - it just means you have more data to review when litigation hits. Third, build a relationship with your legal department before a crisis. The time to figure out how to execute a litigation hold is not the day you receive a preservation notice.
E-discovery is not going to get simpler. The regulatory environment is tightening, data volumes are growing, and courts are becoming less patient with organizations that cannot manage their electronic information. Treating this as purely a legal problem is a mistake. It is an information governance problem, and that puts it squarely in our domain.