Skip to main content

1.5.51

Released 21 September 2026

New​

  • Owners can take a runtime profile without restarting. Organization owners can download a stack dump of every running task (/api/v1/debug/pprof/goroutine?debug=2) and memory or CPU profiles, so a job that seems stuck can be diagnosed where it is waiting instead of being cleared by a restart. Every other role is refused.

  • Revoke a single retained version from the certificate's History tab. Each older version now has a Revoke action that invalidates only that version at the CA while the certificate keeps serving its current one; a revoked version shows a Revoked tag and can no longer be rolled back to. Both revoke dialogs now let you pick the revocation reason, such as key compromise or superseded; a revocation that needs approval shows its version and reason to the reviewer.

Improved​

  • Memory limits are respected before the process runs out. CertAutoPilot now reads the memory limit of its container or systemd service and makes the garbage collector work harder at 80 % of it, so a heavy discovery scan slows down instead of being killed by the kernel; an explicit GOMEMLIMIT is used as-is. Refreshing discovery findings after a scan also needs about half the memory it did.

  • Retries of failed jobs are spread out. A failed job's next attempt now varies by up to 20 % around the usual 30 s, 1 min, 2 min… schedule (still at most 10 minutes), so the many batches of one distribution no longer retry against a recovering device in the same second.

  • Cloudflare is no longer marked experimental. The module is exercised end-to-end against a real zone: create, update with the re-keyed id, adoption by hostname and by pinned id, renewal, rollback, geo restriction and the legacy certificate pack.

  • Discovery no longer reports a Cloudflare-only origin as an outage. When a proxied record's origin refuses connections from outside Cloudflare while the same name serves a certificate at the edge, the origin row is tagged Origin · Cloudflare-only with an error that points at the edge row, instead of a bare "dial timeout". It still counts as a failed probe, raises no finding or notification, and turns green the moment the origin opens up.

  • The discovery source's Endpoints tab is grouped on the server. One row per IP:port shows how many names and distinct certificates sit at that address, the managed/unmanaged/failed split and up to three certificate subjects, and a page boundary no longer splits an address. Opening a row lists its names with the certificate each one serves and when it expires.

Changed behaviour​

Behind Cloudflare or another CDN? Name its client-IP header in the config

The backend no longer reads CF-Connecting-IP or True-Client-IP unless server.remote_ip_headers lists them. By default only X-Real-IP and X-Forwarded-For — the headers the bundled nginx sets itself — identify the client. Previously the CDN headers were trusted first even though nginx forwards them untouched, so any visitor could choose the IP written to the audit log and reset every per-IP rate limit (login, code verification, token refresh, setup, ACME challenge) by sending the header. If a CDN that overwrites such a header really fronts your installation, add it to server.remote_ip_headers after upgrading, or rate limiting and the audit log will see the CDN's addresses instead of your users'.

Operators can now revoke certificates

Revocation used to be limited to project admins, which also meant the Require approval for revocation setting never took effect. Operators can now revoke a whole certificate or a single retained version. With Require approval for revocation on, their revocation becomes a request that a project admin approves in My Requests; with it off, operators revoke directly, like renewals and downloads. If revocation should stay behind an admin in your organization, turn the approval setting on before upgrading.

  • An Exchange deployment confirms the certificate on the server before skipping. When the stored record says a certificate is already current, the run now checks every server it names: the certificate must still be there and, when IIS is one of the services, still on the IIS 443 binding. If someone removed or re-bound it by hand, or the check cannot connect, the run redeploys and says why instead of reporting success. Unchanged runs cost one extra management session, plus one connection per additional server when IIS is bound.

  • Deactivating a user ends their access now. Deactivation revokes every session token the user holds and the token refresh is refused for a deactivated account, so the user is signed out within 15 minutes at most instead of keeping access for up to a week. A login that is half-way through two-factor verification is refused at completion. Reactivating restores login with the same roles; the old sessions stay ended.

  • A user disabled in CertAutoPilot stays disabled even if the directory still accepts them. An LDAP login used to re-enable the account on every successful directory bind; it now refuses the login and leaves the account disabled until an administrator re-enables it.

  • A directory identity never takes over a local account. If the directory has an entry with the same username as a local CertAutoPilot user — including the owner created at setup — that login is refused and the local account keeps its password source and roles. Previously the local account was silently converted to LDAP and its roles handed to the directory user. The server logs a warning naming the username and directory entry.

  • A kubeconfig credential must be self-contained. Saving a kubeconfig that uses exec or auth-provider plugins, tokenFile, or file paths for the client certificate, key or cluster CA is refused, at save and again at use; the refusal names the entries. Use the inline *-data fields or a token. Existing credentials of that shape stop working at their next deploy or health check and must be re-created inline.

  • Bulk renew follows the renewal approval setting. While Require approval for renewal is on, operators can no longer start a bulk renewal — the single renew action raises an approval request per certificate instead; admins and owners bypass as before, with one auto-executed record for the bulk action. Bulk actions also observe the licence certificate limit and are limited to 10 per minute per project.

  • A renewal threshold can no longer be used to skip renewal approval. With Require approval for renewal on, an operator's threshold change that would make a certificate due at the next sweep is refused and pointed at the renew action (admins keep the lever). The threshold must be at least 1 day, the key rotation policy reuse or rotate, and the rotation interval at least 1.

  • An approved request needs a requester who still belongs to the project. Executing or retrying an approved request, and consuming an approved download token, now requires the requester to still hold at least the operator role in that project. The request stays approved and a project admin can still execute it.

  • Certificate policies apply to approval-gated requests too. An operator's issue request is checked against the project policy when it is raised and again when the approved request is executed; an execution-time violation fails the request with the violation codes and moves the certificate to rejected so it can be deleted and resubmitted. Previously the approval gate switched the policy engine off for operators.

  • Private keys written over SSH default to mode 0600. A private_key, combined or pfx path-set file with no explicit mode used to take the server's default permissions — world-readable on most hosts — and a refused permission change on such a file now fails it instead of leaving the key readable. Files are staged in a restricted temporary sibling and renamed into place, so a reader never sees a partial file and an interrupted transfer leaves the previous file intact.

:::caution Existing keys are tightened on the next run A private key deployed by an earlier version without an explicit path-set mode is corrected to 0600 on the next run even if the certificate did not change, and the ActionSet does not run for it. A service running as another user that read the old 0644 key keeps working until it restarts. Set the path-set mode explicitly if the key must stay readable by others. The atomic replace also needs write permission on the directory and gives the file a new owner (the SSH user) unless owner is set; where the directory is not writable — a root-owned directory with a pre-created key file — the module keeps rewriting the file in place, as before, and says so in the job log. :::

  • A target group's default credential is validated. Like a target's own credential, it must exist, belong to the project and be a type the group's module accepts; a mismatched type (for example a Vault token on an SSH group) or an unknown credential is refused on create and update.

  • Downloads and exports leave an audit trace. A certificate download now records its format and whether the private key was included; a download by approved one-time token, the certificate export and the audit-log export — all three were unrecorded before — are now audit events of their own.

  • Custom webhook headers that carry a credential are masked. In a webhook notification channel's summary (and in the audit record of an edit), the value of a header such as Authorization or X-Api-Key is shown masked; other headers stay readable. Saving the edit form with the mask in place keeps the stored value.

  • A discovery record accepts at most 64 ports. Discovery looks for certificates on service ports; a single record listing thousands of ports was a port scanner in disguise. Duplicate ports in one record are refused. The same limit applies to the port list of an AXFR or cloud DNS-provider source. A source saved earlier with more ports can still be renamed, paused or rescheduled — the limit applies once its ports are edited.

  • Wrong two-factor codes now count toward the account lockout. A wrong code used to burn only its session's five attempts; it now also advances the account's failed-login counter, and that counter is reset only when a login fully succeeds, so someone holding a password no longer gets a fresh set of code guesses for every new login attempt.

  • An organization always keeps an active owner. Removing the owner role from the last active owner, deactivating that user or deleting them is now refused. Grant the owner role to another active user first.

Fixed​

  • The Helm chart installs with its defaults. The chart's default backend image pointed at a registry path that was never published, so a stock helm install ended in ImagePullBackOff; it now points at the published image, and the post-install notes describe the real workloads and the HTTPS UI address.

  • A kubeconfig credential can no longer run commands on the CertAutoPilot server. A project admin could save a kubeconfig whose user entry ran a program (exec) or pointed tokenFile at a file on the server, and the Kubernetes client executed it as the CertAutoPilot process on the next credential test, health check, picker call or deploy — enough to read the encryption keys and with them every private key of every project. Such entries are now refused.

  • Windows deployments could be made to run an operator's command. Values that reach PowerShell on IIS, Exchange and generic Windows targets — a binding picker's site name, a distribution override's site or host header, transfer paths — were quoted against the ASCII apostrophe only, while PowerShell also treats the typographic quotation marks (‘ ’ ‚ ‛) as quotes. A value carrying one could end the quoted text and run what followed, on the host and with the credential the administrator configured. All five quote characters are now neutralised.

  • A project admin could rewrite another project's certificate policy. The policy save accepted the policy id from the request body, and the update matched on id and version alone, so an id read from project B (visible to any operator there) could be replayed under project A's route to disable B's policy. The saved policy's identity now always comes from the project in the route, and the update is scoped to that project.

  • The web UI no longer receives session tokens in the login response. The browser signs in with cookie_only and gets its session in the httpOnly cookies only, so a script running in the page cannot read the 7-day refresh token from the response. Scripts and integrations that log in without the flag keep receiving the tokens as before.

  • Standalone installs send session cookies with the Secure flag. The bundled nginx only serves HTTPS, but the backend behind it issued cookies without Secure. The installer now sets it, both in a fresh config.yaml and in the service unit, so existing installs get it on their next upgrade.

  • Failed logins, token refreshes and refresh-token reuse detections are now visible in the audit log. These pre-authentication events were filed under an internal placeholder organization that the audit page and the export never show, so a brute-force attempt left no visible trace. They are now attributed to the organization like every other entry.

  • Spreadsheet formula injection in the CSV exports is closed. A certificate name or owner, or the username typed into a failed login, could begin with =, +, - or @ and run as a formula when the certificate or audit export was opened in a spreadsheet. Such cells are now written as text.

  • A DNS provider API key can no longer end up in an error message. A network error from the Namecheap client quoted the request URL, API key included, and that text reached the certificate's last error, the job log and notifications. Query-string keys are redacted, and every error stored on a certificate, a job or a job log is redacted on the way in.

  • A deploy script cannot print the private key into the job log. ActionSet output on SSH and WinRM targets is redacted before it is logged, including key material re-encoded as hex — joined, line-wrapped (xxd -p), space-, colon- or dash-separated (od, hexdump, openssl … -text, PowerShell) — which the existing PEM and base64 rules did not see.

  • Certificate validation errors keep their detail. The redaction that removes an Authorization header from stored errors also matched the word "authorization" in ordinary ACME diagnostics ("invalid authorization: …") and cut off the CA's explanation. Only header-shaped values are redacted now.

  • A terminally failed job no longer stores a secret quoted in its error. The last attempt of a failed job (and the audit entry written for it) went through a path that did not redact; it now does.

  • Sending several two-factor codes at once no longer gets extra guesses. The five-attempt limit of a code session was checked before the code and counted after it, so parallel requests all saw an unused session. The attempt is now reserved first; at most five codes per session are ever checked.

  • Restarts and upgrades no longer leave running jobs stuck for five minutes. On shutdown the service waited up to 55 seconds for running jobs, but the Helm chart and the systemd unit stop it after 30, so it was killed before it could hand those jobs to another worker and they waited out their five-minute lock. It now waits up to 15 seconds and then hands them over immediately.

  • An LDAP server that stops answering no longer blocks logins. A directory that accepted the connection but never replied held the login request indefinitely. Connecting now times out after 10 seconds and each request after 30, and the login fails with an error instead of waiting.

  • A restart no longer resets the renewal check interval to hourly. The scheduler applied the interval configured under Settings → General only when the setting was changed; after an upgrade, a reboot or any service restart it started on the built-in hourly cadence until an administrator re-saved General settings. The renewal, ARI-refresh and distribution sweeps could therefore run up to an hour late after every restart. The persisted interval — and the domain check interval — are now applied as soon as the scheduler starts.

  • Creating several certificates at once no longer fails with "certificate creation is busy" on a licensed instance. With a certificate cap in force, creates are serialised for a moment so the cap can never be overshot; the request that arrived second used to be rejected immediately instead of waiting for that moment. It now waits up to five seconds, which covers a burst of concurrent creates (bulk actions, scripted imports, the API called in parallel). The error still appears if the lock is genuinely stuck for longer. Unlimited licences were never affected.

  • Two published HTTP/2 advisories in the bundled gRPC client library are closed. Neither was reachable (CertAutoPilot runs no gRPC server), but the shipped binary no longer carries them.

  • A webhook rollback now names the certificate it actually sends. The domains field of a rolled-back webhook delivery (and the .Domains template variable, and the SMTP subject of a certificate without a common name) is read from the older version's own certificate rather than the record's current names. Ordinary deployments are unchanged.

  • Re-deploying a certificate that is already on an Exchange server no longer fails with "already exists". Exchange does not report whether a certificate has its private key over its management endpoint, so every run tried to import the certificate again. A retry after a failed service restart, or a rollback to a version still on the server, now re-enables the existing certificate.

  • Editing a discovery source during a scan no longer disturbs that scan. Saving the source used to write back its whole stored copy, which could leave the source marked as scanning after the scan ended, let a second scan start, or move the next scheduled scan back. An edit now changes only the settings you edited.

  • Changing a user's roles can no longer leave them with none. If the change was interrupted partway, the user could lose every role, which locked an organization out when that user was its only owner.

  • A job that succeeded on a retry no longer shows the failed attempt's error. A job that failed once and went through on a later attempt kept the old message on the Jobs list and its detail page next to a green completed status, which read as a failure nobody could explain. The error is cleared when the job completes; the failed attempt is still in the job's logs.

  • An owner can deactivate or delete another owner again. Those two actions checked the caller's role in a way that never saw the organization role, so every owner was treated as a non-owner and refused with "only an owner can deactivate/delete another owner". Admins remain unable to remove an owner.