DEV API LAB / FIELD GUIDE 04

Build GitHub webhooks that recover cleanly

Verify deliveries, accept work durably, track revisions, and reconcile retries without duplicate side effects.

By DevAPI.com™6 min readDev Tools & SDKs
Build GitHub webhooks that recover cleanly — neon typographic artwork with DevAPI.com™ branding

A repository webhook looks simple at first: receive an event, inspect the change, and perform some useful work. The difficulties appear when the same event is processed twice, a dependency is temporarily unavailable, or a result refers to a revision that has already moved. A dependable integration needs a record of what it received and what it actually completed, not just a route that returns a successful response.

This guide designs a pull-request checklist service as an example. The service reads an approved repository, evaluates a bounded set of project rules, and publishes one traceable result. It is not a live DevAPI.com integration. Use the GitHub Dev API topic for the broader permission model and the following workflow to reason about event handling, recovery, and review.

Choose the smallest useful event surface

Begin with the events and actions the product genuinely needs. A checklist service may care about a pull request being opened or updated, but it does not necessarily need every repository event. A smaller subscription surface reduces irrelevant work and makes it easier to explain why the service acted. Keep that explanation visible in the operational record.

The GitHub webhook best-practices documentation covers measures such as selecting required events, verifying deliveries, and processing work asynchronously. Treat the official delivery guidance as part of the integration contract. Verify the current documentation when implementing the receiver rather than copying assumptions from an old sample application.

Verify before trusting the payload

A publicly reachable endpoint should not assume that every incoming request came from the expected sender. Verify the delivery using the documented signature mechanism and the configured secret. Preserve the raw bytes needed for that verification before middleware changes the request body. Reject a failed verification before the payload can trigger privileged work.

Keep the webhook secret separate from the credential used for subsequent API calls. They serve different purposes: one establishes delivery authenticity, while the other authorizes application access. Neither should appear in logs. When rotating secrets, plan the transition deliberately so an operator can distinguish a configuration error from a genuinely invalid delivery.

Validate the expected event shape

After verification, check the event type, action, repository context, and fields the workflow actually uses. A valid signature does not imply that the event is relevant to the current feature. Unknown actions should produce a controlled outcome rather than falling through into a handler that was written for a different payload shape.

Normalize the minimal information needed to begin work: a delivery identifier, repository identifier, pull-request identifier, revision, and the event’s purpose. Keep the original payload available only according to a limited retention policy. A durable queue message should not carry every field merely because the sender included it.

Acknowledge only after durable acceptance

Keep expensive analysis out of the request handler. The receiver should verify, validate, and durably record or enqueue the work, then respond according to the sender’s delivery requirements. The worker can process the checklist independently. This separation prevents a slow repository fetch from consuming the entire delivery response window.

The word durably matters. Returning success after placing an event in an in-memory list can lose work when the process restarts. Decide what must survive a crash before the request is acknowledged. If acceptance fails, return an appropriate failure rather than pretending the event is safely scheduled. Document who or what can recover that failed delivery.

Make duplicates a normal case

Treat duplicate delivery as an expected input to the system design. Use a durable delivery record and a uniqueness constraint or equivalent control. The receiver can recognize work that has already been accepted, while the worker tracks whether it is pending, running, completed, or awaiting reconciliation. Avoid a single boolean that cannot distinguish those states.

Duplicate detection at intake is not sufficient on its own. A worker can crash after publishing a result but before recording completion. Design the publishing operation so a retry updates or reconciles the intended result instead of creating another unrelated message. Keep the relationship between the delivery, the analysis run, and the published result explicit.

Analyze an exact revision

A branch name can move between the event and the later API read. Record the revision that the checklist is meant to describe and fetch the corresponding content where the platform supports it. Include that revision in the result. A green checklist that silently evaluated an older change can be more misleading than an explicit incomplete result.

Decide how the integration handles superseded work. A newer revision might cancel an older pending run, or the service might finish both while clearly labeling their targets. Do not let completion order decide which result is presented as current. Compare revision and workflow state before publishing a status intended to describe the latest change.

Keep repository content untrusted

A repository can contain instructions, build scripts, or files contributed by people who should not control your integration’s privileged environment. Reading a pull request is not permission to execute arbitrary code from it. A checklist service should begin with static inspection and narrowly scoped operations, not a generic shell execution capability.

If the product later requires code execution, design an independent isolation and credential boundary for that task. Do not make the webhook worker both an unrestricted execution environment and the holder of release credentials. The agent tool-calling guide describes a similar separation between proposed work and the application authority that performs it.

Classify failures before retrying

Not every error improves with repetition. A transient transport problem may justify a bounded retry. A missing permission or an invalid repository identifier usually needs configuration or user action. Record the failure category and stop repeated attempts when they cannot help. Use a backoff policy and an overall work limit appropriate to the integration.

Keep uncertain side effects distinct from known failures. A timeout while reading metadata is different from a timeout after submitting a comment. The latter may require checking whether the comment exists before trying again. Present an honest reconciliation state to operators rather than repeatedly publishing until one response happens to reach the worker.

Create an operator recovery path

Build a way to inspect accepted deliveries and their outcomes without exposing secrets or complete confidential payloads. An operator should be able to find the repository, target revision, current state, retry history, and relevant diagnostic identifier. Those fields make a recovery decision possible without searching unrelated application logs.

Define how failed work can be replayed. A replay should pass through the same authorization and duplicate-protection boundary as ordinary processing. Record that it was requested and why. A manual recovery path is not a license to bypass controls; it is a deliberate way to resume a workflow when the automatic path cannot safely continue.

Test crashes at the inconvenient points

Create tests for a repeated delivery, an irrelevant event, a failed signature, a queue failure, and a missing permission. Then test process interruption after durable acceptance, during analysis, and immediately after publishing. The most valuable tests exercise the gaps between steps, where a happy-path demonstration provides little evidence.

Use controlled fixtures and a staging repository rather than performing experiments against an important production workflow. Check that one logical operation produces one intended visible result after reconciliation. For the command-line side of the operating process, continue with CLI and SSH automation and its distinction between failed, cancelled, and uncertain outcomes.

Conclusion: keep a record of the work

Reliable webhooks are less about making the receiver clever and more about making the workflow explicit. Verify the delivery, accept it durably, target an exact revision, and maintain enough state to recover without repeating side effects blindly. A small checklist integration built this way gives both developers and operators a clear answer to the essential question: what did this event actually cause the system to do?