DoorDash's AI Agents Tackle 60K Feature Flags
Alps Wang
Sep 19, 2026 · 1 views
AI's Pragmatic Leap in Code Maintenance
DoorDash's application of multi-agent LLMs for feature flag cleanup is a compelling demonstration of AI's potential to automate tedious and complex engineering tasks. The system's ability to handle dependency injection patterns, which often complicate automated refactoring, is particularly noteworthy. By integrating live experimentation data, engineer approval, and robust automated validation, they've built a system that not only identifies stale flags but also generates reliable pull requests, significantly reducing manual effort and cost. The reported average cleanup time and cost per flag, when compared to manual efforts, highlight a substantial efficiency gain. This approach is a testament to how LLMs can move beyond simple code generation to sophisticated code transformation and maintenance.
However, the reliance on specific LLMs like Claude Sonnet and Opus, and Google's Agent Development Kit, ties this solution to a particular ecosystem. While the underlying principles are transferable, the direct implementation might require significant adaptation for organizations using different AI platforms or cloud providers. The mention of engineer intervention for complex cases, particularly those involving deep call chains and cross-interface parameter threading, indicates that full automation is still a future goal. The one-hour timeout per agent, while a necessary safeguard, might also extend the overall cleanup process for highly complex scenarios. Future work on confidence scoring and post-cleanup code quality checks will be crucial for broader adoption and trust.
Key Points
- DoorDash has developed a multi-agent LLM system to automate the cleanup of over 1,000 stale feature flags.
- The system leverages live experimentation data, engineer approval, isolated Git worktrees, and automated validation.
- It significantly reduces the time and cost of manual feature flag cleanup, with an average of 13.8 minutes and $4.79 per cleanup.
- The approach is designed to handle complex dependency injection patterns where flag logic is distributed across multiple files.
- The workflow involves an orchestrator agent and specialized cleanup agents operating concurrently.
- Validation checks include builds, tests, code coverage, and static analysis before generating a pull request.
- While successful for simple and medium complexity flags, complex scenarios still require engineer intervention.
- Future plans include confidence scoring and post-cleanup code quality passes.

📖 Source: DoorDash Uses Multi Agent LLMs to Clean up 60,000 Feature Flags
Related Articles
Comments (0)
No comments yet. Be the first to comment!
