DATA FUNDAMENTALS.
The SOV is only one document in a property submission. The forms and reports around it carry just as much critical data — and just as many ways for that data to get stuck. Here’s the unstructured data problem, and what it takes to pull clean data out of ACORD forms and loss runs.
In an earlier post we took apart the Statement of Value — the spreadsheet at the heart of every property submission. But it never travels alone. It arrives bundled with ACORD forms, loss runs, and supplemental applications. Each carries data the underwriter needs, and each comes in its own format.
Together, these documents are where the industry’s “unstructured data problem” actually lives. The information is all there — it’s just trapped in document layouts and inconsistent formats your systems can’t read directly. Let’s look at the two biggest culprits, then at what it takes to get usable data out of each.
WHAT IS AN ACORD FORM?
ACORD — the Association for Cooperative Operations Research and Development — is the non-profit that maintains the standardized forms used across the insurance industry. For commercial property, the workhorses are forms like the ACORD 125 (Commercial Insurance Application) and the ACORD 140 (Property Section), which capture applicant details, premises and building information, coverage requests, and remarks.
These forms almost always arrive as PDFs. That’s the good news. An ACORD PDF is a consistent, well-defined document, and modern extraction can read it directly, with no re-keying. The catch is that a PDF form, however tidy it looks, still isn’t structured data. The values sit inside a page layout, not in queryable fields, and the details that decide a risk are often tucked into free-text remarks.
WHAT IS A LOSS RUN?
A loss run is the claims history of a risk — a report that lists every claim over a period (usually three to five years) with its date, cause, status, amounts paid, reserves, and total incurred. It’s how an underwriter separates a clean risk from one that quietly bleeds water-damage claims every winter.
Loss runs have no shared industry layout — every carrier formats its own. But the format they arrive in matters enormously for whether the data can be used. A loss run delivered as a spreadsheet is machine-readable: its rows, columns, and figures parse cleanly into structured claims data. The same report flattened into a static document is much harder to work with reliably. If you want loss history to flow automatically into your systems, an Excel loss run is what gets you there.

A PDF form is digital, but it isn’t structured data. The values are visible to a person and invisible to your rating system — until something turns the layout back into fields.
THE UNSTRUCTURED DATA PROBLEM, DEFINED
Here’s the core issue. Underwriting systems, pricing models, and catastrophe platforms all need structured data: clean fields, consistent units, predictable columns. But submissions arrive as documents built for human eyes — PDF forms, varied spreadsheets, and free-text notes. Every submission needs someone, or something, to bridge the gap between a document and a dataset.
Two things make the gap hard to cross. First, format sprawl: even a “standardized” ACORD form exists in countless filled-in variations, and loss runs have no shared layout at all, so no two carriers’ reports line up. Second, meaning in the margins: the decisive details often live in remarks fields, footnotes, and one-off notes that don’t map neatly to any column. Traditional automation handles the easy, predictable cases and hands everything else to a person, who then spends the day re-keying figures, hunting for values across tabs and attachments, and reconciling formats before the actual underwriting can begin.
The bottleneck was never the underwriting. It was getting the data into a shape you could underwrite from.
WHY IT MATTERS
Unstructured intake is a hidden tax on the whole operation. It slows quote turnaround, so brokers take their business to whoever answers first. It introduces re-keying errors that flow straight into pricing. And it caps capacity: a team can only process as many submissions as it can manually decode, so good risks get set aside for no reason other than a stubborn document.
CLOSING THE GAP
The answer isn’t to fight the formats — brokers will keep sending PDFs and spreadsheets in every shape imaginable. The answer is to turn them into data automatically. Ping reads ACORD forms directly from PDF, extracts SOV data from spreadsheets, and parses Excel loss runs into structured claims data. We validate and enrich as we go, and deliver the results straight into your systems in minutes.
That means the ACORD remarks box gets captured, the Excel loss run becomes clean claims data, and missing fields get flagged or filled — automatically. Your team starts from decision-ready data instead of a stack of documents to decode.
Turn the documents in every submission into data.
See how Ping extracts clean, structered data from ACORD PDFs, SOV spreadsheets, and Excel loss runs.