10 min read
Garbage In, Optimized Out? Why Trusted Operational Data Matters More Than Big Data

Written by Dr. Ümit Kuvvetli
Founder & Chief Optimization Scientist
Cities already have the data. That isn't the problem. The problem is knowing which data deserves to be trusted.
Data doesn't create decisions.
Trusted data does.
Every day, public transport agencies collect millions of GPS positions, smart card transactions, GTFS schedules, passenger counts, vehicle diagnostics, and traffic records. Together, these datasets promise a complete picture of how a transport system performs. But operational reality is rarely that simple. More data does not automatically produce better decisions—sometimes it only produces greater confidence in the wrong conclusion.
The Assumption Behind Every Intelligent System
Artificial intelligence, optimization, digital twins, simulation, and predictive analytics all rely on the same silent assumption: the operational data accurately represents what actually happened.
When that assumption fails, every downstream decision becomes less reliable. The algorithm is rarely the problem. The evidence is.

When Correct Data Leads to Wrong Decisions
Imagine an AVL system reporting that a bus remained at one stop for more than three minutes. An optimization model interprets excessive dwell time; schedulers add recovery time; operations managers investigate delays. Everything appears rational—until someone visits the location. A wheelchair passenger was boarding. The operation performed exactly as intended. The coordinates were accurate. The interpretation was not.
Now imagine another situation. GPS shows a vehicle travelling several kilometres without carrying passengers. Reports identify excessive dead mileage; fuel efficiency appears poor; management questions performance. Field verification tells a different story: the vehicle had completed its final trip and was returning to the depot. Again, the data was correct. The operational context was missing.
Or consider a fare validation. A passenger taps the card twice within seconds. Without validation logic, analytics records two boarding events. In reality, the first transaction simply failed. One passenger, one journey, two records—and one misleading conclusion.
Data doesn't describe reality.
It samples reality.
Sensors never observe operations; they observe measurable events. GPS records positions. Validators record transactions. GTFS publishes planned service. Passenger counters estimate occupancy. None of these systems truly understands what planners need to know: Was the service reliable? Was the passenger experience acceptable? Did operations perform as intended?
Reality begins where raw data ends.

The Hidden Cost of Trusting Bad Evidence
Poor-quality operational data does far more than distort dashboards. It quietly reshapes decisions. An incorrect timestamp changes travel-time estimation; travel-time estimation changes running times; running times change vehicle scheduling; scheduling changes fleet requirements; fleet requirements change budgets. One unreliable observation can influence hundreds of operational decisions.
Optimization does not remove uncertainty. It often amplifies it.
Validation Is Operational Thinking
Validation is often confused with cleaning—removing duplicates, fixing formats, deleting null values. Those tasks improve databases. They do not necessarily improve decisions.
Operational validation asks different questions. Does this GPS position make physical sense? Is this fare validation consistent with the vehicle's location? Does this travel pattern represent passenger behaviour—or sensor failure? Those questions cannot be answered by SQL alone. They require operational understanding, and often a second look at GTFS data quality against service reality.
What Research Continues to Demonstrate
Academic research increasingly reaches the same conclusion. A large-scale study of Fortaleza's public transport system integrated AVL, GTFS and fare collection data before any behavioural modelling was attempted. Researchers first performed extraction, normalization, compatibility analysis, spatial matching and extensive validation before searching for patterns. Only then did meaningful analysis begin.
Their conclusion is difficult to ignore: exploratory analysis should be preceded by rigorous data treatment—yet many studies continue to overlook this step.
Even more revealing is what happened to the data itself. Although more than 330,000 passenger records were initially available, only around 118,000 were considered sufficiently reliable for analysis. Nearly two-thirds of the available data was intentionally excluded—not because it was unusable, but because it could not support confident decisions.
Sometimes better decisions begin with less data.
Validation Requires Multiple Perspectives
No operational dataset is complete on its own. AVL explains movement. Smart cards explain demand. GTFS explains intention. Traffic explains disruption. Field observations explain everything the other systems cannot.
Modern transport research consistently concludes that reliable decision-making emerges by validating independent datasets against one another rather than trusting any single source of truth. Confidence is built through agreement—not volume.
Why Field Intelligence Still Matters
No algorithm understands operational context the way an experienced planner does. Field intelligence explains what GPS cannot: why passengers avoid a perfectly accessible stop; why vehicles consistently bunch along one corridor; why actual operations diverge from perfectly designed schedules; why two identical datasets can produce completely different planning decisions—a theme that also appears in the human factor of public transport optimization.
Operational reality still begins on the street, not inside the database.

The OW Perspective
This philosophy shapes every analytical workflow we build. Before optimization, OW validates. Before simulation, OW reconciles. Before prediction, OW questions the evidence.
That is why modules such as OW GTFSHub™ and OW OD Matrix™ begin with operational consistency rather than optimization itself. Improving a system without first understanding whether its evidence is trustworthy only accelerates the wrong decisions. Related decision layers—OW FreqOpt™ for headways and OW CostLogic™ for cost attribution—only become decision-grade after that foundation holds.
Our objective is not simply to process more information. It is to increase confidence in every decision built upon it.
Better Decisions Begin Earlier
Artificial intelligence will continue to improve. Optimization models will become more sophisticated. Digital twins will become increasingly realistic. But none of those technologies can compensate for evidence that fails to represent operational reality.
The organizations that lead the next generation of public transport will not be those collecting the largest datasets. They will be the ones building the most trusted ones.
Better algorithms cannot rescue unreliable evidence.
They can only optimize it.
And the future of public transport will not be built on more data. It will be built on trusted operational data.
Related Modules
Used in this piece
Related Posts
Continue with adjacent topics—from mixed-integer programming (MIP) and combinatorial optimization to multi-objective scenario modeling in public transit.

The Passenger Who Never Boarded: Transit Planning's Invisible Problem
Denied passengers and latent demand: why non-boarders never appear in ridership data. Lessons from Eskişehir, Lisbon, and Abu Dhabi on measuring real transit demand.

What People Actually Expect from Public Transport (And Why It's Still Not Being Met)
IPPR research applied to Turkey: passengers expect reliability, affordability, and safety. How OW GTFSHub™, CostLogic™, and FreqOpt™ help municipalities meet those expectations.




