Engineering teams generate enormous volumes of simulation and test data. Most of it is never used again. It sits in proprietary solver formats on HPC storage, on individual workstations, and inside PLM systems. Opening it usually means going back to the original tool. Reusing it to train AI is harder still.
We built the Simr Data Platform to change that. It turns proprietary CAE (computer-aided engineering) simulation results and physical test data into structured, queryable, versioned datasets. Datasets that are ML-ready, meaning curated, labeled, accurate, and carrying their engineering context. This is a behind-the-scenes look at how we built it, and at how it fits into an engineering team's day.
Start with the data itself. A single crash, CFD (computational fluid dynamics), FEA (finite element analysis), thermal, or acoustics run can produce gigabytes of output. That output is written in solver-specific formats. Each format has its own structure. Each one usually needs its own tool to read.
Getting one value out is harder than it sounds. Say you want the peak pressure on one surface, or the temperature at one node over time. Today that often means writing a script for that specific solver and that specific file. Every engineer writes their own. The scripts are rarely shared. The knowledge stays in one person's head.
The data is also scattered. Simulation results sit in one place. Lab test measurements sit in another. There is no single index you can query across both. So most engineering data is used once, for one decision, and then forgotten.
This is a real problem for AI. Surrogate models and reduced order models (fast approximations that stand in for a full physics simulation) need large, clean, labeled datasets to learn from. Most companies cannot assemble those datasets from what they already have. We call this the cold-start problem for Engineering AI. You want to train models on your own physics. You cannot, because your own physics is locked inside files you cannot easily read.
Solving the data problem was only the first step. The harder part was fitting the solution into how engineers actually work. Three principles guided that.
First, the data does not move. The Simr Data Platform runs inside the customer's own cloud account, or on their on-premise servers. Engineering data never leaves the customer's environment. For teams working on unreleased products, that is not a nice-to-have. It is a requirement.
Second, the platform does the tedious work automatically. It parses and analyzes simulation results files on its own. There is no manual setup for each new file or format.
Third, engineers stay in control. The Visual Query Editor lets an engineer select the parts and sections of a results file they care about, and they do it visually. No scripting needed. Point at the surface, the region, or the time step you want. The platform handles the extraction.
From there, the platform produces outputs for decision making. 3D, 2D, and 1D visualizations. Reports. KPI tables. All generated automatically from the curated data.
The data is also made available to large language models through MCP (Model Context Protocol, an open standard for connecting AI models to data sources). An engineer can ask questions about their own results in plain English. Not through a proprietary interface. Your data, in your own language.
For teams building their own models, the Simr Data Platform API gives full programmatic access to that curated data. The data is labeled, accurate, and contextual. It feeds AI training, surrogate modeling, reduced order modeling, design optimization, and other advanced analytics techniques.
A platform like this only earns its place if it works on real problems, across different physics and different kinds of data. Our early deployments cover exactly that range.
At a Fortune 500 company that designs popular wearables, the Simr Data Platform automates the analysis of audio measurements coming out of the test labs. This is physical test data, not simulation.
At an automotive customer, it automates the data pipeline from crash safety simulations into the customer's in-house AI models. Results flow from the solver, through curation, and into training, without manual handling at each step.
At a leading autonomous driving technology company, it is used in the thermal analysis of sensors.
Three domains. Acoustics, crash safety, and thermal. Two kinds of data. Physical test and simulation. One platform handling all of it.
This is where we are today. We are just getting started. If your team is sitting on simulation and test data you cannot reuse, we would like to talk.
Simr builds engineering data infrastructure for Physics AI. The Simr Data Platform turns proprietary simulation and test data into structured, queryable, versioned, ML-ready datasets, running inside your own cloud account or on your on-premise servers so your data never leaves.