Ir al contenido principal
Operaciones

DNSSEC validation failures during our move off Cloudflare

Why Cloudflare was added in 2024

Since roughly 19 August, some DNSSEC-validating resolvers have intermittently been unable to resolve onetimesecret.com and its subdomains. The failure could return SERVFAIL, depending on which authoritative server a resolver had contacted. We did not announce this maintenance in advance, and we only learned the scope from an external report on 26 August. That is on us.

The issue occurred during our migration of authoritative DNS from Cloudflare to Bunny. For a period, both providers were in the .com delegation and signing the zone with different keys without a shared DNSKEY record set. As a result, a resolver that had cached keys from one provider could fail to validate data signed by the other. Cloudflare was added reluctantly during the September 2024 DDoS, and this migration is the long-planned removal of that dependency.

Custom domains were unaffected. They are served through onetime.co, which never used Cloudflare and did not experience these DNSSEC validation failures.

Mitigation is now in place: Bunny's ZSK (18608) was added to Cloudflare's DNSKEY set, stale Hetzner nameserver records were removed, and the Cloudflare zone was deleted. Cloudflare can continue to serve the deleted zone until it is purged. We still need to remove Cloudflare's nameservers and DS record (2371) from the parent delegation at Porkbun (we're not able to remove CF's nameservers without also deleting the DS records). We will not consider this incident resolved until that work is complete and DNSSEC validation has been verified from independent resolvers.

A temporary DNS probe now monitors the remaining cleanup. It queries independent DNSSEC-validating resolvers over DNS-over-HTTPS. For onetimesecret.com and catch.onetimesecret.com, it requires successful DNSSEC validation and an A-record answer. For CDN-fronted hostnames, where the upstream CDN zone is unsigned, it verifies that an A-record answer is still reachable. This is temporary monitoring until a more robust solution is in place, but it will alert us if the remaining delegation work causes the same type of failure again.