https://grafana.com/blog/2024/02/21/how-the-open-source-caddy-server-uses-grafana-cloud-for-full-stack-observability/ Path: blog/ 2024-02-08-how-caddy-server-leverages-its-sponsored-grafana-cloud-pro-account-for-cloud-native-full-stack-observability.list-en.md Copy path to clipboard Copied! * Products Open source Solutions Learn Docs Company * Downloads Contact us Sign in Create free account Contact us Products All Products Core LGTM Stack Logs powered by Grafana Loki Grafana for visualization Traces powered by Grafana Tempo Metrics powered by Grafana Mimir and Prometheus extend observability Performance & load testing powered by Grafana k6 Continuous profiling powered by Grafana Pyroscope Plugins Connect Grafana to data sources, apps, and more end-to-end solutions Application Observability Monitor application performance Frontend Observability Gain real user monitoring insights Incident Response & Management with Grafana Alerting, Grafana Incident, Grafana OnCall, and Grafana SLO Deploy The Stack Grafana Cloud Fully managed Grafana Enterprise Self-managed Pricing Hint: It starts at FREE Open Source All Open Source Grafana Loki Multi-tenant log aggregation system Grafana Query, visualize, and alert on data Grafana Tempo High-scale distributed tracing backend Grafana Mimir Scalable and performant metrics backend Grafana OnCall On-call management Grafana Pyroscope Scalable continuous profiling backend Grafana Beyla eBPF auto-instrumentation Grafana Faro Frontend application observability web SDK Grafana Agent Batteries-included telemetry collector Grafana k6 Load testing for engineering teams Prometheus Monitor Kubernetes and cloud native OpenTelemetry Instrument and collect telemetry data Graphite Scalable monitoring for time series data Community resources Dashboard templates Try out and share prebuilt visualizations Prometheus exporters Get your metrics into Prometheus quickly Solutions All end-to-end solutions Opinionated solutions that help you get there easier and faster Kubernetes Monitoring Get K8s health, performance, and cost monitoring from cluster to container Application Observability Monitor application performance Frontend Observability Gain real user monitoring insights Incident Response & Management Detect and respond to incidents with a simplified workflow monitor infrastructure Out-of-the-box KPIs, dashboards, and alerts for observability Linux Windows Docker Postgres MySQL AWS Kafka Jenkins RabbitMQ MongoDB visualize any data Instantly connect all your data sources to Grafana MongoDB AppDynamics Oracle GitLab Jira Salesforce Splunk Datadog New Relic Snowflake Learn All Learn Stay up to date GrafanaCON 2024 Our biggest community event New - Reg ObservabilityCON on the Road 2024 Open source observability conference Story of Grafana 10 years of Grafana Observability Survey 2023 Key findings and results Blog News, releases, cool stories, and more Events Upcoming in-person and virtual events Success stories By use case, product, and industry Technical learning Documentation All the docs Webinars and videos Demos, webinars, and feature tours Tutorials Step-by-step guides Workshops Free, in-person or online Writers' Toolkit Contribute to technical documentation provided by Grafana Labs Plugin development Visit the Grafana developer portal for tools and resources for extending Grafana with plugins. new Join the community Community Join the Grafana community new Community forums Ask the community for help Community Slack Real-time engagement Grafana Champions Contribute to the community new Community organizers Host local meetups new Docs All Docs Grafana Cloud Grafana Grafana Agent Grafana Loki Grafana Mimir Grafana Tempo Grafana Pyroscope Grafana OnCall Application Observability Grafana Faro Grafana Beyla Grafana k6 OpenTelemetry Prometheus Grafana Enterprise Grafana Enterprise Logs Grafana Enterprise Traces Grafana Enterprise Metrics Enterprise plugins Community plugins Grafana Alerting Visit Docs Get started Get Started with Grafana Build your first dashboard Getting started with Grafana Cloud What's new / Release notes Company All Company Our team Careers We're hiring Events Partnerships Newsroom Contact us Merch Help build the future of open source observability software Open positions Check out the open source projects we support Downloads Sign in Core LGTM Stack Logs powered by Grafana Loki Grafana for visualization Traces powered by Grafana Tempo Metrics powered by Grafana Mimir and Prometheus extend observability Performance & load testing powered by Grafana k6 Continuous profiling powered by Grafana Pyroscope Plugins Connect Grafana to data sources, apps, and more end-to-end solutions Application Observability Monitor application performance Frontend Observability Gain real user monitoring insights Incident Response & Management with Grafana Alerting, Grafana Incident, Grafana OnCall, and Grafana SLO Deploy The Stack Grafana Cloud Fully managed Grafana Enterprise Self-managed Pricing Hint: It starts at FREE Grafana Cloud Free forever plan (Surprise: it's actually useful) * Grafana, of course * 14 day retention * 10k series Prometheus metrics * 500 VUh k6 testing * 50 GB logs, traces, and profiles * and more cool stuff Create free account No credit card needed, ever. Grafana Loki Multi-tenant log aggregation system Grafana Query, visualize, and alert on data Grafana Tempo High-scale distributed tracing backend Grafana Mimir Scalable and performant metrics backend Grafana OnCall On-call management Grafana Pyroscope Scalable continuous profiling backend Grafana Beyla eBPF auto-instrumentation Grafana Faro Frontend application observability web SDK Grafana Agent Batteries-included telemetry collector Grafana k6 Load testing for engineering teams Prometheus Monitor Kubernetes and cloud native OpenTelemetry Instrument and collect telemetry data Graphite Scalable monitoring for time series data Community resources Dashboard templates Try out and share prebuilt visualizations Prometheus exporters Get your metrics into Prometheus quickly end-to-end solutions Opinionated solutions that help you get there easier and faster Kubernetes Monitoring Get K8s health, performance, and cost monitoring from cluster to container Application Observability Monitor application performance Frontend Observability Gain real user monitoring insights Incident Response & Management Detect and respond to incidents with a simplified workflow monitor infrastructure Out-of-the-box KPIs, dashboards, and alerts for observability Linux Windows Docker Postgres MySQL AWS Kafka Jenkins RabbitMQ MongoDB visualize any data Instantly connect all your data sources to Grafana MongoDB AppDynamics Oracle GitLab Jira Salesforce Splunk Datadog New Relic Snowflake All monitoring and visualization solutions Stay up to date GrafanaCON 2024 Our biggest community event New - Reg ObservabilityCON on the Road 2024 Open source observability conference Story of Grafana 10 years of Grafana Observability Survey 2023 Key findings and results Blog News, releases, cool stories, and more Events Upcoming in-person and virtual events Success stories By use case, product, and industry Technical learning Documentation All the docs Webinars and videos Demos, webinars, and feature tours Tutorials Step-by-step guides Workshops Free, in-person or online Writers' Toolkit Contribute to technical documentation provided by Grafana Labs Plugin development Visit the Grafana developer portal for tools and resources for extending Grafana with plugins. new Join the community Community Join the Grafana community new Community forums Ask the community for help Community Slack Real-time engagement Grafana Champions Contribute to the community new Community organizers Host local meetups new Featured Getting started with the Grafana LGTM Stack We'll demo how to get started using the LGTM Stack: Loki for logs, Grafana for visualization, Tempo for traces, and Mimir for metrics. Watch now - Grafana Cloud Grafana Grafana Agent Grafana Loki Grafana Mimir Grafana Tempo Grafana Pyroscope Grafana OnCall Application Observability Grafana Faro Grafana Beyla Grafana k6 OpenTelemetry Prometheus Grafana Enterprise Grafana Enterprise Logs Grafana Enterprise Traces Grafana Enterprise Metrics Enterprise plugins Community plugins Grafana Alerting Visit Docs Get started Get Started with Grafana Build your first dashboard Getting started with Grafana Cloud What's new / Release notes Grafana: 10.3 Grafana Mimir: 2.9 Loki: 2.9 Tempo: 2.3 Our team Careers We're hiring Events Partnerships Newsroom Contact us Merch Grot cannot remember your choice unless you click the consent notice at the bottom. Site search Ask Grot - AI Beta (what could go wrong?) [ ] I am Grot. Ask me anything Grot good Grot bad Feedback Related resources Blog post Docs OSS project I'm a beta, not like one of those pretty fighting fish, but like an early test version. So, our vampires, I mean lawyers want you to know that I may get answers wrong. - Go back Feedback Write a short description about your experience with Grot, our AI Beta. Rate your experience (required) Comments (required) [ ] Send Sending... Sent Thank you! Your message has been received! [ ] Blog * All * Community * Culture * Engineering * News * Release [ ] Blog Menu [ ] * All * Community * Culture * Engineering * News * Release How the open source Caddy server uses Grafana Cloud for full-stack observability Mohammed Al Sahaf Mohammed Al Sahaf * February 21, 2024 * 5 min --------------------------------------------------------------------- Mohammed Al Sahaf serves as Technical Product Manager at Samsung Electronics Saudi Arabia. Outside his day job, he serves with the Caddy team to tackle the web of problems facing web servers in the third millennium. Mohammed is the author of Kadeessh, formerly caddy-ssh, and the maintainer of numerous Caddy modules. When he isn't programming, he is trying to catch up on life and sleep with the help of coffee. You can find his caffeinated wonders at caffeinatedwonders.com. As maintainers of the OSS project Caddy server, we know the value of end-to-end observability, in terms of accelerating root cause analysis and reducing MTTR. But rather than do all the undifferentiated heavy-lifting to manage an observability stack, my team wants to dedicate as much time as possible to our core mission: developing and optimizing our open source project. These were some of the biggest reasons behind our recent migration to Grafana Cloud. In this blog post, I'll take a closer look at why the Caddy team chose Grafana Cloud as an observability solution, and how it'll help us advance the Caddy project, moving forward. Get a free Grafana Cloud Pro account for your OSS project Full-stack observability is essential for any OSS project to progress and thrive. That's why Grafana Labs -- a company with open source software in its DNA -- offers a free Grafana Cloud Pro account to maintainers of OSS projects like Caddy. Interested in receiving your Cloud Pro account? You can reach us at community@grafana.com. Our path to Grafana Cloud Caddy, an extensible web server written in Go, is known as the first generally available HTTP/2 server. Caddy's flagship features include automatic HTTPS/TLS procurement and management through the ACME protocol, as well as robust OCSP stapling. When Caddy launched in 2015, users could create a one-of-a-kind, on-demand Caddy build, using a custom set of plugins through the download page. This custom builder withstood the test of time, as demand for Caddy grew. The "buildworker," as the Caddy dev team calls it, was rewritten shortly after the GA of Caddy v2 in 2020. The rewrite was to accommodate a new module structure and build flow using xcaddy, a convenient library and tool to generate custom Caddy builds, under the hood. The new version had been functioning with minimal complaints -- until early 2023. At this point, we started to hear about unresponsive download requests and the download page failing to build the requested custom Caddy build. We were relying on our cloud provider's dashboard to monitor server health, and would manually comb through journald to identify and troubleshoot issues, as needed. The cloud provider's dashboard provided a high-level view of server resource utilization, but didn't offer deep insights into the applications running on the server. The absence of a unified and user-friendly interface for log inspection also made it cumbersome to troubleshoot when away from the keyboard, to share interesting log lines with team members, and to trace problematic requests across service boundaries. To make matters worse, this was reactive: end users needed to tell us, so we couldn't act on or resolve issues proactively, due to a lack of alerting. Recognizing these growing pains, we knew we needed a more mature observability solution for our infrastructure. So in March 2023, we connected with Grafana Labs, and learned they have a program through which they offer a free Grafana Cloud Pro account to OSS project maintainers (a special shout-out here to Dave Henderson, a fellow member of the Caddy project and a senior software engineer at Grafana Labs, who helped us discover this program). This Grafana Cloud Pro account enabled several Caddy project maintainers to monitor our infrastructure and, after experimenting with the dashboards provided by the Linux Server integration, we found value almost immediately. Faster root cause analysis with Grafana Cloud To get started with Grafana Cloud, we followed the integration onboarding steps to install and configure Grafana Agent on all our servers. Knowing our buildworker was the bottleneck, we customized a dashboard to inspect various angles of the buildworker node. The node_exporter was vital to understand how our workload affected the instance it's running on. A screenshot of a Caddy dashboard in Grafana. We quickly learned we had a goroutine leak due to long-running builds that included some of our heavier modules. We observed in a time series chart the number of goroutines growing over long periods, but never touching the zero line: A time series chart showing the number of goroutines growing over time. Long-running builds had not been a problem until recently, which we knew may be a symptom. Timeouts were set on build requests, so we knew at least one trigger for the struggling server. Treating this symptom ensured jobs that were stuck didn't run indefinitely or consume resources. When we fixed the goroutine leak on the buildworker, we saw that there may be spikes, but the chart line was not always stepping up: A time series chart that shows a resolved goroutine leak. Although we addressed the goroutine leak, we still needed to find and resolve the root cause of our server performance issue. The question remained: why were those builds slow? Compilation demands RAM, CPU, and I/O, depending on the stage. We observed reports of unresponsiveness that never correlated with excessive RAM utilization, but CPU utilization had always been high -- to the point of failing to accept SSH logins within reasonable time. This was a key realization for us to discover that our process is CPU-bound, not memory-bound. It provided us with a new action plan: the buildworker server upgrade should optimize for CPU. What's next for Caddy and Grafana Cloud While our plan was a success, we are far from finished. SRE is a process of continual improvement. Looking ahead, the observability tools in Grafana Cloud will allow us to efficiently diagnose and treat issues as they arise. Log aggregation allows us to correlate various events across different application and system boundaries; continuous monitoring lets us alert on sudden drops in activity and the events surrounding the drops; profiling indicates where resources utilization is subpar; and synthetic monitoring provides an external perspective to observe our system as a black box, using periodic PING and/or TRACEROUTE. We look forward to advancing our observability strategy -- and the Caddy OSS project -- with Grafana Cloud. To learn more about a Grafana Cloud Pro account for your OSS project, reach out at community@grafana.com. Tags Engineering On this page * Get a free Grafana Cloud Pro account for your OSS project * Our path to Grafana Cloud * Faster root cause analysis with Grafana Cloud * What's next for Caddy and Grafana Cloud Scroll for more Up next Wei Li * 7 Feb 2024 * 5 min read New in Grafana k6: The latest OSS features in v0.49.0 and static IPs in Grafana Cloud k6 Grafana k6 v0.49.0 is here, featuring a built-in web dashboard for real-time result visualization and tons of other improvements. k6 Engineering Read more Ryan Perry * 6 Feb 2024 * 5 min read Combining tracing and profiling for enhanced observability: Introducing Span Profiles Span Profiles, a new feature in Grafana Cloud and Grafana OSS, enables deeper analysis of both tracing and profiling data for a more granular view of... Continuous profiling Engineering Read more Sriramajeyam Sugumaran * 5 Feb 2024 * 3 min read Infinity plugin for Grafana: Grafana Labs will now maintain the versatile data source plugin It's an exciting new chapter for the Infinity data source plugin, which lets you seamlessly visualize data from JSON, CSV, XML, and GraphQL endpoints... Plugins Engineering Read more Sign up for Grafana stack updates [ ] [ ] [ ] Subscribe Sorry, an error occurred. Email update@grafana.com for help. Note: By signing up, you agree to be emailed related product-level information. --------------------------------------------------------------------- --------------------------------------------------------------------- * Grafana * Overview * Deployment options * Plugins * Dashboards * Products * Grafana Cloud * Grafana Cloud Status * Grafana Enterprise Stack * Grafana Cloud Application Observability * Grafana Cloud Frontend Observability * Grafana Cloud IRM * Grafana Cloud k6 * Grafana Cloud Logs * Grafana Cloud Metrics * Grafana Cloud Profiles * Grafana Cloud Traces * Grafana SLO * Open Source * Grafana * Grafana Loki * Grafana Mimir * Grafana OnCall * Grafana Tempo * Grafana Agent * Grafana k6 * Prometheus * Grafana Faro * Grafana Pyroscope * Grafana Beyla * OpenTelemetry * Grafana Tanka * Graphite * GitHub * Learn * Grafana Labs blog * Technical documentation * Downloads * Community * Community forums * Community Slack * Grafana Champions * Community organizers * Grafana ObservabilityCON * GrafanaCON 2024 * The Golden Grot Awards * Successes * Workshops * Videos * OSS vs Cloud * Load testing * Company * * The team * Press * Careers * Events * Partnerships * Contact * Getting help * Merch --------------------------------------------------------------------- Grafana Cloud Status Sitemap Legal and Security Terms of Service Privacy Policy Trademark Policy Copyright 2024 (c) Grafana Labs Grafana Labs uses cookies for the normal operation of this website. Learn more. Got it!