AlertD launched out of stealth, unveiling its agentic AI SRE (Site Reliability Engineering) and DevOps platform designed to tackle the mounting operational complexity of cloud operations.
AlertD empowers teams to quickly get contextualized visibility across their AWS environments through an easy to use user interface that delivers verifiable data and allows for powerful collaboration.
AlertD was founded in 2024 by Geoff Hendrey (former Cisco Distinguished Engineer, AppDynamics Chief Architect, Splunk Principal Architect) and Freddy Mangum (former Cisco Entrepreneur-in-Residence, Fortinet VP of Products & Marketing, Venture Capital Cybersecurity Advisor) after experiencing firsthand the limitations of legacy observability, log analysis, cloud ops and alerting tools.
"I have decades of experience working with some of the most powerful observability tools in the industry," explains Geoff Hendrey, co-founder and CEO of AlertD. "While these legacy tools delivered rich instrumentation, they still required extensive manual setup to configure alerts that might serve as early indicators of production issues. But as application development velocity has accelerated—and with the rise of complex microservices architectures—SRE and DevOps teams are now struggling to keep up with the scale and demands of maintaining production uptime."
Hendrey's experience with foundational technologies—including work that preceded what we now know as Retrieval-Augmented Generation (RAG)—combined with years of enterprise experience at Cisco, AppDynamics, and Splunk, made it clear that thoughtfully applied LLMs have the potential to fundamentally transform the lives of SRE and DevOps teams.
The name 'AlertD' pays homage to the Unix daemon ('d')—the silent processes that power critical infrastructure. The company's vision is to build a suite of specialized AI SRE and DevOps agents that are always available, always helpful, and always working in the background to support uptime-critical teams.
The founding team brings deep expertise in launching and scaling products and scaling them from pre-revenue to hundreds of millions in revenue. AlertD has raised $3 million in pre-seed funding, led by Puneet Agarwal of True Ventures, emphasizing capital efficiency and hands-on partnership.
Agarwal is well known for backing visionary founders early, having led the first investments in companies like Duo Security (acquired by Cisco for $2.35B), Puppet Labs (acquired by Perforce), and numerous other foundational infrastructure startups.
"Geoff and Freddy bring a rare combination of technical depth and go-to-market instinct, shaped by decades of experience with some of the most widely used observability, monitoring, analysis and cybersecurity platforms in the industry," said Puneet Agarwal, partner at True Ventures. "With AlertD, they're applying LLMs in a way that has real potential to change how SRE and DevOps teams work day to day, bringing clarity and speed to some of the most demanding moments in software operations."
AlertD's AI SRE and DevOps platform has been shaped directly through ongoing collaboration with mid- to large-sized enterprises that operate mature SRE and DevOps functions. These design partnerships have enabled the team to focus on delivering measurable proactive and reactive outcomes—streamlining daily workflows for individual contributors while providing leadership with real-time visibility into system health and team performance.
Ryan Raines, Sr. Director of DevOps at Privateer—the geospatial intelligence and space sustainability company founded by Steve Wozniak in 2023—explained: "With 19 years in the industry, I lead one of the most talented SRE and DevOps teams out there. Yet even with great people, the demands we face have outpaced what we can solve by simply adding headcount.
"SREs and DevOps spend nearly 50% of their time on low-value work—not due to inefficiency, but because today's tooling for managing production uptime is overly complex while our environments continue to scale. I want my people focused on high-value work, which is why we're helping shape what an optimal AI-native tool for SREs and DevOps should look like.
"The time to embrace AI agents in SRE and DevOps is now. Just as developers have adopted AI co-pilots to accelerate coding, we must adopt intelligent automation to improve uptime. Despite the noise in the AI tooling space, few solutions truly address the breadth of needs our teams face. That's why we partnered with AlertD—their deep expertise and vision for transforming SRE and DevOps workflows makes this a tool our team will rely on daily for both proactive and reactive operations."
AlertD is not just another AI debugger relegated to incident response—it's a comprehensive and extensible platform designed for the full spectrum of cloud operations. AlertD's AI agentic platform is purpose-built to operate seamlessly within the AWS ecosystem, while maintaining a cloud-, LLM-, and SDLC-agnostic architecture. The platform supports both proactive and reactive use cases, empowering SRE and DevOps professionals to work alongside their existing tools and gain comprehensive insights across security, compliance, cost optimization, troubleshooting, account ownership, and infrastructure management. Through its intuitive interface and natural language capabilities, AlertD democratizes access to AWS environment data, enables verification of information retrieved by AI agents, and delivers actionable insights that help teams save significant time and costs.
"Today's cloud operations are overwhelmed by noise and manual toil. The volume and velocity of work demands and ticket queue outpace human capacity to drive fast, effective outcomes," said Freddy Mangum, Co-founder and COO of AlertD. "SRE and DevOps professionals are highly skilled, yet too often trapped in reactive workflows that limit their impact."
Mangum continued, "Specialized AI SRE and DevOps agents are the natural next step—operating 24/7 on behalf of cloud operations teams to filter noise, synthesize complex data across systems, and surface actionable, contextual insights within seconds. While many incumbents are introducing AI agents, most remain to reactive use cases.
"In contrast, AlertD is a multi-purpose, multi-environment–agnostic platform designed to complement and extend existing AI solutions. Just as GitHub Copilot, Cursor, or Claude enhances developer productivity during the build phase, AlertD empowers teams during the run and production phase. Our vision is to make AlertD the 'Slack for production uptime'—an indispensable tool that gives cloud operations teams the confidence and clarity to manage infrastructure at scale."
Key Differentiators:
- AI SRE and DevOps Agents – Specialized AI agents designed for specific proactive and reactive operational tasks.
- Natural Language Interface – Query complex environments and receive actionable insights using plain language—no need for specialized syntax or scripts.
- Proactive and Reactive Operations – Surface insights from AWS metrics and resources to identify and act on security, cost, compliance, and optimization opportunities.
- SRE and DevOps Expert Driven Innovation – Co-built and validated with real-world enterprise design partners to ensure practical, scalable use cases.
- Cloud Security – Deploy securely within your own AWS Virtual Private Cloud (VPC) to maintain full data ownership and control.
- Toil Relief – Obtain AWS insights in seconds—eliminating hours of manual scripting or context switching across multiple tools.
- LLM Agnostic – Vendor-neutral architecture supporting OpenAI, Anthropic, Meta, and other leading LLM providers.
- AI Transparency – Full visibility into AI reasoning, data sources, and analyses, empowering users to verify and trust results.
- Team Collaboration – Easily share knowledge, queries, and insights generated by AlertD AI agents across teams.
- Powerful Search – Advanced search capabilities enable users to drill into AWS metrics and resources for deeper analysis and actionability.
The Latest
Production incidents rarely announce themselves as database problems. They appear as slow transactions, timeouts, rising response times, or an application struggling under a workload it previously handled. APM provides an essential starting point. It can identify a slow transaction path, highlight an affected service, and show that a database dependency is consuming more time than expected. But identifying the database as part of the problem is not the same as explaining what is happening inside it ...
Cloud teams are under constant pressure to reduce spend without slowing development or increasing operational risk. They are deploying autoscalers, rightsizing workloads, enforcing resource requests, reviewing utilization dashboards, and building FinOps processes around cloud-native environments. Yet the results often disappoint ...
Ask most IT leaders about their biggest concern with AI and you'll hear the same answer: hallucinations ... Today, however, the conversation has shifted ... As organizations move beyond chatbots and experiments, they are increasingly deploying AI agents that perform multi-step tasks. These systems retrieve documents, query databases, call APIs, generate reports, write code, and make recommendations. The issue is not whether the model can reason. The issue is whether the organization can see, verify, and govern the decisions being made along the way ...
While organizations want to take control of their telemetry, building telemetry pipelines from scratch can be a very daunting, complicated task, even when leveraging open-source standards like OpenTelemetry. It requires specialized knowledge across distributed systems, data engineering, and security. This fragmented approach across systems causes higher operational costs; it puts a strain on resources and reduces efficiency as teams have to work with different interfaces and processes ...
For decades, enterprise networks were designed around a simple assumption: work happened inside the office. Applications lived in centralized data centers, employees connected through internal infrastructure, and security focused on protecting the perimeter that surrounded everything ... But the way organizations operate today bears little resemblance to that environment. Cloud platforms host critical applications, employees connect from homes and airports as often as they do from offices, and partners collaborate through shared systems that exist far beyond corporate walls. In short, the corporate network no longer resembles the environment it was designed to protect ...
As an analyst who researches how IT organizations design, build, and operate their networks, I find that network data is a constant source of pain. Network teams struggle with data quality, fragmentation, authority, access, and trust. And these issues undermine everything they try to do. Here are the numbers: Only 45% of network teams are completely confident in the accuracy of their network source of truth, which documents the intent of their network ...
The 2026 Global Data Center Survey from Uptime Institute reveals an industry navigating workforce constraints, escalating outage expenses, even as rising costs remain the top concern for management teams ...
The next observability gap may not be in the code. It may be under the rack. That sounds strange until you think about how AI incidents actually feel in the middle of an investigation ... The application dashboard may be accurate. It may also be stopping at the wrong boundary. AI systems depend on software, but they also depend on a dense physical stack: racks, power paths, thermal margin, maintenance activity and, in many environments, liquid cooling. Those physical dependencies can change slowly before they look like a software incident ...
Certificate expiration is the rare outage you can see coming. Every TLS certificate carries the date it stops working, so the moment it will begin breaking connections is knowable in advance. That's what makes an expired certificate such a frustrating way to lose a service. What's changing now is how often that date comes around ...
Enterprises operate different combinations of workloads across cloud, hybrid and multicloud environments. For business-critical workloads, teams need to consider monitoring and observability early so they can detect health issues, investigate failures, and understand operational impact. Organizations place workloads on cloud platforms based on a combination of technical requirements, economics, existing dependencies, organizational standards, and business priorities. Their monitoring priorities therefore depend on what they operate and where those systems run. Those priorities will not look the same for every organization ...