1
The Problem That Wouldn’t Go Away
When I took over leadership of our metrics team, one of my first tasks was reviewing our Azure DevOps data refresh pipelines. That’s when I stumbled upon something that made me do a double-take. Our test execution metrics pipeline was taking 14 hours to complete for 143 projects. Fourteen hours. That meant we could only refresh our metrics once a day, and teams were making decisions based on stale data.
I sat down with the team to understand what was happening. The answer I got was uncomfortable but all too familiar in software engineering. This was legacy code that had been left untouched because the team was afraid of breaking something during a code migration. The original system was designed for just two projects. As the organization grew and more projects were added, nobody revisited the architecture. The processing time just kept growing, and everyone accepted it as “just how things are.”
But here’s the thing about technical debt. It doesn’t stay static. It compounds. What worked for two projects was now choking on 143 projects, and the team had normalized a fundamentally broken system.
Digging Into the Mess
I pulled up the code for a review, and what I found was a masterclass in how not to handle hierarchical data. The pipeline was using five nested for loops to traverse the Azure Test Plan folder structure. It would search for the PI level, then iterate through sprints, then look for regression suites, then manual suites, and finally automation suites. Each level required traversing the entire hierarchy again.
But here’s what really got me. The Azure DevOps API was already returning all the folder hierarchy data in the first API call. The entire tree structure, parent-child relationships, everything we needed was right there in the initial response. Yet our code was making repeated API calls at each level, fetching the same data over and over, and then searching through it linearly to find what we needed.
Every time we wanted to find a child node, we’d iterate through the entire dataset again. For 143 projects with deeply nested test plan hierarchies, this meant thousands of unnecessary iterations and API calls. The algorithm’s time complexity was effectively O(n × m × p × q × r) where each variable represented a level in the hierarchy. No wonder it took 14 hours.
The Solution: Build It Once, Query It Fast
The fix required rethinking how we stored and accessed the hierarchy data. Instead of repeatedly traversing flat lists, I designed a hash-tree structure that would let us look up parent-child relationships in constant time.
The core idea was simple but powerful. We’d make one API call to get all the folder data, then build a dictionary where each node’s ID mapped directly to its complete information including parent references and a set of children IDs. Here’s the function that made it happen:
def build_tree(data: dict) -> dict: construct = {}
for i in data: if i['id'] not in construct: construct[i['id']] = { 'parent_id': int(i['parent']['id']) if 'parent' in i else None, 'id': i['id'], 'name': i['name'], 'children': set() } else: construct[i['id']]['parent_id'] = int(i['parent']['id']) if 'parent' in i else None construct[i['id']]['name'] = i['name']
if 'parent' in i: parent_id = int(i['parent']['id']) if parent_id in construct: construct[parent_id]['children'].add(i['id']) else: construct[parent_id] = { 'parent_id': None, 'id': i['parent']['id'], 'name': i['parent']['name'], 'children': {i['id']} }
return constructThis single pass through the data gave us a complete tree structure where finding any node’s children or parent was an O(1) operation. No more nested loops. No more linear searches. Just direct dictionary lookups.
The second piece of the puzzle was addressing the API rate limiting. We implemented parallel processing with four worker threads. This was deliberate. Azure DevOps API has rate limits, and we found that four parallel workers gave us the optimal balance between speed and staying under the rate limit threshold. Too few threads and we’d leave performance on the table. Too many and we’d hit rate limits and actually slow down.
The Results Were Immediate
The impact was impossible to miss. The pipeline that used to take 14 hours now completed in 30 minutes. That’s a 96% reduction in processing time. But the wins went beyond just speed.
Our API call volume dropped dramatically. Instead of making redundant calls at each level of the hierarchy to fetch data we already had, we were now making targeted calls only to the specific test suites we needed results from. We’d build the tree structure once, identify the leaf nodes we cared about, and fetch only their execution data.
This meant our metrics could now refresh multiple times per day instead of once. Teams had access to near-real-time test execution data. Sprint planning became more accurate. Bug triage meetings had current information. The downstream effects rippled through the entire engineering organization.
What Made This Work
Looking back, there were a few key insights that made this optimization successful. First, I had to challenge the assumption that the existing system was untouchable. Legacy code isn’t sacred just because it’s old or because people are afraid to change it. Sometimes the best way to prevent data loss is to fix the underlying problem rather than working around it.
Second, the solution wasn’t about adding complexity. It was about using the right data structure for the job. Nested loops work fine for small datasets, but hierarchical data needs hierarchical structures. A hash-tree gave us the O(1) lookups we needed without adding significant memory overhead.
Third, we didn’t over-optimize. Four parallel threads was enough. We could have pushed for more, but that would have meant dealing with rate limiting, retry logic, and exponential backoff. Sometimes good enough is actually perfect.
Lessons for Your Own Codebase
If you’re dealing with performance issues in your own systems, here’s what I’d suggest looking for. Start by questioning assumptions, especially around code that “nobody touches.” Often these systems are slow not because they were designed poorly for their original use case, but because the use case changed and nobody revisited the architecture.
Look at your data structures. If you’re doing repeated linear searches through collections, there’s probably a better way. Hash maps, trees, and sets exist for a reason. Use them.
Pay attention to API usage patterns. Are you making redundant calls? Are you fetching data you already have? Sometimes the biggest wins come from eliminating work rather than making work faster.
And finally, measure everything. We knew exactly how long the pipeline took before and after. We could see the API call reduction. Metrics gave us confidence that the changes worked and justification for the time spent optimizing.
The best optimizations don’t just make things faster. They make systems more maintainable, more scalable, and more pleasant to work with. That’s what transforming a 14-hour nightmare into a 30-minute task really means.