How a vehicle manufacturer cut data processing time 15x

  • 15x Faster processing
  • 37 minutes Versus 9 hours for daily geocoding
  • 100 million Records processed each day
Overview

Processing 100 million records a day—faster than ever before

One of the world's largest commercial vehicle manufacturers operates a massive connected fleet, generating nearly 100 million sensor records every day. To manage this data, the company had built a purpose-built, 24-node on-premises Spark Hadoop cluster. But as data volumes and analytic ambitions grew, the infrastructure struggled to keep pace—leaving critical workloads running for hours and dozens of high-value AI and machine learning (ML) use cases out of reach.

Challenge

A Spark infrastructure that couldn't scale to meet growing AI demands

The company's existing Spark Hadoop environment was engineered specifically for its workloads, yet performance remained a persistent bottleneck. A daily geocoding job required 9 hours to complete. A sensor normalization incremental load across hundreds of parameters took 90 minutes. And a full initial load of all historical sensor data demanded up to 36 hours of processing time.

Beyond the raw performance constraints, the infrastructure simply could not support the 20+ new high-value AI/ML use cases the business wanted to pursue—leaving significant analytical potential untapped. The team needed a faster, more cost-effective path forward, and they needed it proven quickly.

Programmers working with multiple monitors displaying code
Solution

An AI-accelerated migration to Teradata—proven in weeks

Teradata's forward-deployed engineering team was given just a few weeks to demonstrate they could run the same workloads faster, cheaper, and at scale. The approach: convert thousands of lines of PySpark code into SQL and Teradata user-defined functions (UDFs), unlocking the full power of Teradata's industry-leading massively parallel processing (MPP) engine—deployed on a small, equivalently sized cluster in Microsoft Azure.

The code conversion was largely automated using AI coding tools that translated PySpark logic directly into SQL. For portions of the code that resisted straightforward translation—particularly inefficient Python UDFs—the team deployed a dedicated AI coding skill enabling an AI agent to generate highly efficient native Teradata UDFs in a compiled language, eliminating the overhead of interpreted Python entirely.

The result: a repeatable, AI-accelerated migration path that delivers 10x Teradata performance gains without requiring engineers to manually rewrite thousands of lines of code.

overhead view spanning bridge with cars moving on highway
Outcome

Dramatically faster workloads—and 20+ new AI use cases now within reach

The performance improvements were immediate and dramatic. The daily geocoding workload dropped from 9 hours to just 37 minutes. Sensor normalization went from 90 minutes to 6 minutes. And the full historical data load was reduced from up to 36 hours to just 3 hours—on an equivalently sized system.

Compared to the company's existing Spark environment, Teradata delivered up to 15x better performance. Against a Databricks implementation evaluated in parallel, Teradata outperformed by 6x to 10x. Beyond raw speed, the benchmark confirmed that 20+ additional high-value AI/ML use cases—previously unachievable on the existing infrastructure—could now be delivered cost-effectively with equivalent resources. For a company processing close to 100 million records per day, that means dramatically more intelligence from the same investment.

Learn more

Let’s connect

Find out how Teradata helps you accelerate business outcomes and provide the business agility you need.

Our sales representatives are here to help.



I consent that Teradata Corporation, as provider of this website, may occasionally send me Teradata Marketing Communications emails with information regarding products, data analytics, and event and webinar invitations. I understand that I may unsubscribe at any time by following the unsubscribe link at the bottom of any email I receive.

Your privacy is important. Your personal information will be collected, stored, and processed in accordance with the Teradata Global Privacy Statement.