An ESP32 that downloads a binary over HTTPS and reboots has demonstrated remote installation. It has not yet demonstrated a production update system.
Production OTA must answer a longer chain of questions. Is this artifact authentic? Is it compatible with this hardware, bootloader, partition layout, configuration, and current security state? Can the device survive interruption? What proves that the new image is healthy? When may it become permanent? What stops the next cohort if failure rates rise? How does an operator know which devices never completed the transition?
These are different decisions, and treating them as one boolean called update successful makes failures harder to contain. ESP-IDF provides useful device-side mechanisms, but a fleet still needs a release contract around them.
The map below separates the device state machine from the fleet decision that allows a release to reach more devices. Failure branches are explicit because retry, rejection, rollback, and rollout stop require different owners and actions.
Confirmation protects the recovery path of one device. It is evidence for cohort promotion, not permission to promote the cohort automatically.
Define the release contract before the API calls
Start by defining an update as a state transition with an owner. The release record should identify the artifact, version, build provenance, supported hardware revisions, minimum compatible bootloader and configuration, security version, intended cohort, health criteria, stop conditions, and recovery action.
The device also needs enough immutable or protected identity to decide whether the artifact applies to it. A URL alone is not an eligibility rule. At minimum, the decision normally depends on device model or board revision, current firmware, partition capacity, and a product-specific compatibility policy. If optional peripherals or regional radio settings change behavior, those facts may also belong in eligibility.
Keep three outcomes distinct:
- installed means bytes were written and the candidate was selected;
- booted means execution reached the new application;
- confirmed means the application produced enough evidence to become the accepted version.
A fourth outcome belongs to the fleet controller: rollout accepted means evidence from the current cohort is strong enough to expose more devices. One healthy device cannot make that decision for the fleet.
Design partition capacity as a recovery budget
ESP-IDF's safe application-update mode requires at least two OTA application slots and an OTA Data partition. The new image is written to the application slot that is not currently selected. After verification, the OTA selection data is updated so that the bootloader can choose the candidate on the next restart. The OTA Data partition itself uses two sectors to reduce the risk that a power failure while changing boot selection leaves no usable selection record.
This is the foundation of an A/B-style application update: do not overwrite the only known bootable image while downloading its replacement. It does not mean every ESP32 product needs the same CSV. A factory application is optional in some layouts, flash sizes differ, signed images add overhead, and application growth competes with data partitions.
Choose the layout from an explicit budget:
- current and expected future application size, including signing and alignment overhead;
- two usable OTA slots if application rollback is required;
- persistent configuration, calibration, logs, certificates, and queued data;
- recovery or factory behavior, if the product uses it;
- migration space required when stored data changes format.
Application rollback does not automatically roll back NVS, filesystems, calibration, or cloud-side schema. If the trial image performs a destructive data migration before it is confirmed, the previous image may boot successfully and still be unable to understand the new state. Prefer backward-compatible or staged migrations, record their version, and test rollback across the boundary.
Power-loss testing must cover more than the binary download. Interrupt the process while erasing, writing, finalizing, changing boot selection, booting the candidate, and migrating persistent state. The expected result is not always immediate success; it is a known bootable and diagnosable state.
Authenticate the artifact, not only the connection
HTTPS protects a transfer channel when certificate validation and trust configuration are correct. It does not by itself prove that the firmware artifact passed your release process, and a checksum detects changed bytes without proving who authorized them.
ESP32 offers several controls that solve different problems:
- signed application verification ties an image to a trusted signing key;
- Secure Boot v2 verifies authorized bootloader and application images before execution on supported chips;
- flash encryption protects selected contents in off-chip flash and is separate from Secure Boot;
- anti-rollback can reject an application whose security version is lower than the value recorded on the device.
ESP-IDF can verify signed OTA applications without hardware Secure Boot, which may be useful in some designs, but that choice is not equivalent to a complete verified boot chain. Conversely, Secure Boot does not operate the release process for you. Private-key custody, signing authorization, key rotation, emergency response, build provenance, and access to the update service remain outside the signature check.
Treat eFuse changes as provisioning decisions, not tutorial toggles. Secure Boot, flash-encryption release mode, and anti-rollback affect recovery and service options. The exact sequence must be tested on the actual ESP32 revision and manufacturing flow using Espressif's current documentation. A generic article cannot safely substitute for that plan.
Use trial boot as a real health gate
With application rollback enabled, ESP-IDF can label the next image as new and move it to ESP_OTA_IMG_PENDING_VERIFY on its first boot. The application then either calls esp_ota_mark_app_valid_cancel_rollback() to accept the image or marks it invalid and reboots with esp_ota_mark_app_invalid_rollback_and_reboot(). If the candidate resets or loses power before confirmation, the bootloader can mark it aborted and return to the previous working OTA application.
The mechanism is useful only if confirmation means something. Calling the valid function at the first line of app_main() proves little more than instruction execution. Define a short, bounded self-test for the product. Depending on the device, it may include:
- expected hardware revision and partition schema;
- readable required configuration and calibration;
- successful initialization of critical peripherals;
- a valid credential and usable system time policy;
- watchdog, storage, and memory behavior inside known limits;
- ability to enter the product's safe operating mode;
- a minimal communication or local-control check when connectivity is part of the operating promise.
Do not make confirmation depend on a service that may legitimately be unavailable for hours unless the device can remain safely pending for that period. ESP-IDF documents a single first-boot confirmation model when rollback is enabled: an unexpected reset before confirmation can trigger a revert. The health test therefore needs to be fast enough to complete reliably and meaningful enough to reject a broken release.
Not every defect appears during startup. A memory leak, radio regression, battery drain, or malformed telemetry may take longer. Use the boot self-test to protect the device's immediate recovery path, then use cohort observation to decide whether the fleet release proceeds. Device confirmation and fleet promotion are two gates.
Treat anti-rollback as a security policy with a cost
Anti-rollback prevents a device from accepting firmware below its recorded security version. That can block reintroduction of a known vulnerable release, but it also narrows recovery choices. On the original ESP32, the security version is represented through a limited eFuse field, and programmed eFuse bits cannot be cleared. ESP-IDF also documents restrictions around factory and test partitions in the anti-rollback scheme.
Do not increment the security version for every ordinary build. Define when a release closes a security boundary that must never be crossed backward, confirm the candidate before advancing the device's irreversible state, and retain at least one compatible recovery image at the permitted security version. Test what happens when an older artifact is offered, when a candidate fails health checks, and when no remaining image satisfies the policy.
Version names and security versions are different namespaces. A semantic version helps humans reason about product releases; a monotonic security version enforces a minimum acceptable boundary on the device. Document the mapping and ownership of both.
Roll out in cohorts with explicit stop conditions
ESP-IDF governs a device's boot choice. It does not decide which thousand devices should update on Tuesday. The fleet service needs its own release state and must tolerate devices that are offline, low on power, behind a captive network, or already in recovery.
A practical rollout can move through internal devices, a small representative canary group, one or more wider cohorts, and finally general availability. Representative matters more than merely small: include hardware revisions, network types, power conditions, regions, and usage patterns that can expose different failure modes.
Each cohort needs a time window and stop conditions defined before release. Useful signals include:
- download, verification, installation, first-boot, confirmation, and rollback counts;
- time spent in each state and devices that stop reporting during a transition;
- reset reasons, watchdog events, boot loops, and storage failures;
- connectivity, queue depth, power, thermal, and application-health regressions;
- distribution by hardware revision and previous firmware, not only a fleet-wide average.
A controller should be able to pause new assignments without preventing already healthy devices from operating. It should distinguish retryable delivery failures from rejected artifacts and unhealthy candidates. Repeated automatic retry of an incompatible image is load, not recovery.
Fleet rollback is also not simply a command called downgrade. Anti-rollback rules, data migrations, backend compatibility, and devices that never installed the candidate may produce several valid versions at once. Define which versions the backend supports during the rollout window and test mixed-version behavior deliberately.
Make update state observable and bounded
Record enough state to reconstruct what happened without streaming sensitive logs indefinitely. A useful device update report can contain device identity, hardware revision, current and target version, artifact identifier, update state, attempt count, timestamps or monotonic durations, last error class, boot reason, confirmation result, and rollback result.
Use stable error classes for fleet decisions, while retaining detailed local diagnostics where appropriate. NETWORK_UNAVAILABLE, ARTIFACT_REJECTED, INSUFFICIENT_SPACE, WRITE_FAILED, HEALTH_CHECK_FAILED, and ROLLED_BACK lead to different actions. A single OTA_FAILED counter does not.
Reports may be duplicated or delayed after reconnect, so state transitions should be idempotent and ordered with a device-local sequence or comparable mechanism. Bound retained history and redact secrets, URLs carrying credentials, certificates, and raw memory. Observability should support a decision, not create a second uncontrolled data product.
Inventory completes the picture. Operators must be able to enumerate devices by current version, target version, cohort, hardware revision, update state, and last contact. Devices that are indefinitely absent require a support policy: wait, require physical service, revoke, or retire. Silence is a state to manage, not a successful rollout.
Test the failure paths before the first remote release
A production test plan should exercise transitions, not only the happy-path demo. Include at least:
- loss of power or forced reset during download, flash write, selection update, first boot, and confirmation;
- truncated, corrupted, incorrectly signed, wrong-hardware, oversized, and security-version-rejected artifacts;
- full or corrupted persistent storage and incompatible configuration;
- failure of each critical peripheral used by the startup health test;
- unavailable update service, expired or invalid transport trust, slow links, and interrupted resume;
- rollback after a data-schema change and operation with mixed firmware versions;
- a fleet cohort that crosses a stop threshold while other devices are offline.
Record the expected state after each interruption and the evidence that proves recovery. Repeat the tests on production-equivalent flash, partition tables, security configuration, and hardware revisions. A development board with permissive security settings cannot validate the manufacturing image's recovery behavior.
A production OTA readiness checklist
Before enabling remote updates for a pilot, verify that:
- every artifact has provenance, compatibility metadata, a signature policy, and an owner;
- the partition layout preserves a tested bootable application during update;
- persistent-data migrations remain compatible with rollback or have their own recovery plan;
- the new image stays provisional until a bounded product-specific self-test passes;
- signing keys and release authorization are separated from the public update service;
- Secure Boot, flash encryption, and anti-rollback decisions are documented independently;
- device update states and errors are visible without exposing secrets;
- cohort assignment, pause, stop conditions, and mixed-version support are implemented;
- power loss, invalid artifacts, failed health checks, and rollback have been rehearsed;
- a named person or team can stop the rollout and coordinate recovery.
ESP-IDF supplies strong primitives for writing an inactive slot, choosing the next boot image, verifying software, confirming a trial application, and reverting an unhealthy one. Production quality comes from connecting those primitives to an artifact policy, useful health evidence, fleet-level gates, and a recovery procedure that has already failed safely in testing.
The goal is not to make every update succeed. It is to prevent one failed release from becoming an uncontrolled fleet-wide state.