The wrong way to choose between MQTT and HTTP is to compare two packet captures from a clean laboratory connection and declare the protocol with fewer bytes the IoT winner. A production device pays for much more than headers. It pays for connection setup, keepalive traffic, radio wake time, retries, offline storage, authentication, backend state, operational tooling, and recovery when acknowledgement does not mean what the application assumed.
MQTT and HTTP also describe different interaction models. MQTT is a client-server publish/subscribe transport: clients publish messages to topics and subscribe to topic filters through a broker. HTTP is organized around requests to target resources, with method and response semantics understood by clients, servers, gateways, and caches.
The useful question is not which protocol is better. It is which interaction model makes the product's state transitions easier to define, observe, retry, and operate.
The map below starts with the workload rather than a protocol logo. Its three outcomes are candidates for validation, not automatic architecture decisions.
The bottom rail matters most: identity, bounded state, safe retry, versioning, and observability remain product responsibilities whichever branch is selected.
Start with the workload, not the device label
Calling a product an IoT device says almost nothing about its communication pattern. A battery sensor that wakes twice a day to upload one batch, a mains-powered gateway streaming events, and an actuator waiting for time-sensitive commands have different needs even if all three use the same microcontroller.
Describe the workload before selecting a protocol:
- Who initiates communication: device, service, or both?
- Is each exchange a named resource operation, an event, a stream of observations, or a command?
- Does one message need one consumer or many independent subscribers?
- How long is the device online, and may it keep a connection open?
- What must survive an outage, and where is that state stored?
- Can duplicate delivery occur, and what makes processing idempotent?
- How quickly must a server-originated command reach a connected device?
- Which broker, gateway, identity, logs, limits, and on-call procedures can the team actually operate?
This description often makes the answer visible before implementation details do.
A practical decision matrix
Telemetry
- MQTT tendency: many small live events flow through a broker to multiple consumers.
- HTTP tendency: the device uploads bounded records or batches to one service.
- Verify: batch size, reconnect cost, and the loss policy.
Commands
- MQTT tendency: a connected client receives broker-delivered server-originated messages.
- HTTP tendency: the device polls, fetches desired state, or receives commands through another channel.
- Verify: required latency, offline expiry, and authorization.
Fan-out
- MQTT tendency: publishers should not need to know every consumer.
- HTTP tendency: one receiving service owns distribution to downstream consumers.
- Verify: subscriber isolation, slow consumers, and backpressure.
Resource operations
- MQTT tendency: topic events are the natural domain model.
- HTTP tendency: GET, PUT, POST, DELETE, conditions, and status codes express the operation clearly.
- Verify: retry and concurrency semantics.
Intermittent connectivity
- MQTT tendency: broker session state and subscriptions add useful continuity.
- HTTP tendency: wake-request-sleep and a device-owned queue are simpler.
- Verify: session expiry, queue bounds, and the real radio duty cycle.
Existing platform
- MQTT tendency: the team can operate broker identity, ACLs, sessions, and topic governance.
- HTTP tendency: the team already operates gateways, APIs, tracing, and rate limits.
- Verify: total operational ownership, not only firmware effort.
The matrix is directional, not absolute. MQTT 5 can model request/response with response topics and correlation data. HTTP can support long-lived connections, streaming, and server push patterns through additional technologies. The question is how much application convention is required before the chosen protocol resembles the interaction model you actually need.
Where MQTT earns its place
MQTT is strongest when publish/subscribe is part of the product architecture rather than a transport optimization. A publisher can emit an observation without knowing which analytics, alerting, storage, or control services consume it. The broker matches topics to subscriptions and performs fan-out. That decoupling can keep device firmware independent from changes in downstream consumers.
A long-lived connection also gives the broker a path for server-originated messages to a connected device. Commands, desired-state notifications, or configuration events can arrive without device polling. The Last Will mechanism can notify subscribers after an abnormal connection loss, although it should be interpreted as a connection/session signal, not proof that hardware is broken.
MQTT sessions can preserve subscriptions and protocol state across reconnects when configured with a non-zero Session Expiry Interval. For intermittently connected devices, that can continue delivery of eligible QoS messages after reconnection. Retained messages can provide the last retained value to a new matching subscriber.
Those features need boundaries. A retained message is not a history. A broker session is not an unlimited device-local queue. A Will is not a complete presence system. Session expiry, message expiry, broker limits, subscription design, and what happens after state is discarded must be explicit product decisions.
MQTT is a good candidate when several of these are true:
- the device is frequently connected and receives asynchronous events;
- publish/subscribe fan-out is fundamental to the backend;
- topic-based routing and independent consumers reduce coupling;
- session continuity and delivery state are useful during short outages;
- the team can govern topic namespaces, per-device authorization, quotas, retained state, and broker capacity.
Where HTTP is the clearer tool
HTTP fits naturally when the device acts on named resources or performs bounded exchanges. Fetch the current configuration. Upload a batch of measurements. Report the result of a job. Request a signed firmware manifest. Download an artifact with range or cache semantics. Replace desired state conditionally. These operations have clear request, response, status, and resource boundaries.
HTTP method semantics also help reason about retries. RFC 9110 defines safe methods and identifies PUT, DELETE, and safe methods as idempotent: repeating the same intended request has the same intended server effect as applying it once. That does not make every application safe automatically, but it provides a vocabulary for designing retryable operations.
POST is not inherently idempotent. If a device may repeat an upload after losing the response, give the operation a stable event or request identifier and make the service deduplicate it, or expose a resource whose identifier lets the device use an idempotent PUT. Status codes alone cannot tell a client whether an earlier connection failure happened before or after the server committed a side effect.
HTTP should not be dismissed as one connection per reading. HTTP/1.1 defaults to persistent connections, and ESP-IDF's HTTP client can reuse a connection when multiple requests use the same handle and the server leaves it open. HTTP/2 adds multiplexed streams where the stack and device resources support it. A short-lived battery device may still choose to connect, send a bounded batch, receive a result, and sleep instead of maintaining a broker session.
HTTP is a good candidate when several of these are true:
- exchanges are naturally resource-oriented request/response operations;
- uploads are sparse or can be batched;
- the device mainly initiates communication;
- ordinary API gateways, status codes, conditional requests, tracing, and rate limits are valuable;
- direct server-to-device delivery is unnecessary or can use polling, desired-state fetches, or a separate channel.
QoS and status codes do not prove business success
MQTT defines three delivery levels. QoS 0 is at most once. QoS 1 is at least once and duplicates can occur. QoS 2 uses a multi-step acknowledgement flow to deliver exactly once within the MQTT sender/receiver protocol exchange.
That last phrase matters. A QoS acknowledgement does not prove that a database transaction committed, an actuator moved, a downstream subscriber completed processing, or a bridge did not create a second delivery in another system. Application side effects still need identifiers, deduplication, state checks, and result reporting appropriate to the consequence. QoS 2 also adds protocol state and exchanges, so it should solve a measured requirement rather than serve as a decorative maximum setting.
HTTP has a similar boundary. A 200 or 204 response can represent application success according to the API contract, while a transport timeout leaves the client uncertain whether the server applied the operation. Retries are safe only when the method and application semantics make them safe.
For either protocol, define what each acknowledgement proves:
1. bytes accepted by a protocol peer;
2. message admitted by broker or API;
3. application validated the payload;
4. state transition committed;
5. physical action completed;
6. result observed and reported.
Many systems need more than one of these signals. Do not compress them into a single delivered flag.
Offline behavior belongs to the product
Neither protocol removes the need for a bounded local policy. When a device cannot reach the broker or HTTP service, decide which records are retained, how they are ordered, when they expire, and what is dropped when storage fills. Record event time separately from transmission time. Give retryable events stable identity.
With MQTT, decide whether publications are queued in RAM or persistent storage before the connection exists, which QoS is used after reconnect, how Session Present changes behavior, and how broker-side session expiry interacts with device state. With HTTP, decide how batches are formed, whether partial acceptance is possible, how response loss is handled, and how a device resumes a large transfer.
Reconnect policy matters at fleet scale. If every unit retries on the same schedule after an outage, either protocol can create a recovery storm. Use exponential backoff, jitter, admission controls, and server feedback. Bound attempts and expose queue pressure so that operators can distinguish network absence from a device that is permanently incompatible.
Compare complete connection lifecycles
Protocol overhead is workload-dependent. MQTT has compact control packets and can be efficient for repeated small messages over an established connection. That connection also uses memory, broker state, keepalive traffic, and radio time, and it must reconnect when NAT, access networks, power saving, or server policy closes it.
HTTP carries method, target, headers, and response metadata, but connection reuse changes the cost substantially. Batching may amortize both protocol and TLS overhead. A device that wakes rarely may spend less total energy completing one bounded request than maintaining liveness, even if the individual HTTP exchange contains more bytes. Another device sending events every second may favor a long-lived messaging connection.
Measure the complete duty cycle on representative hardware and networks:
- DNS, TCP, and TLS setup frequency;
- bytes and radio-on time during useful transfer;
- keepalive, polling, and idle timeout behavior;
- reconnects under packet loss and address changes;
- RAM, flash, and persistent-queue writes;
- server state and cost per connected or active device;
- recovery load after a shared outage.
A header-size table cannot answer these questions.
Security depends on identity and authorization, not the protocol name
Both protocols can use TLS. Neither is secure merely because its URI starts with a secure scheme or a library enables encryption. The device must validate the server trust chain and name, protect its credential, rotate or revoke identity, constrain interfaces, and avoid logging secrets. The service must authorize each device for the smallest required action.
MQTT authorization usually needs topic-level rules for publish and subscribe. Wildcards, retained messages, shared subscriptions, and administrative topics can expand authority unexpectedly if the namespace is treated as naming decoration. A device that may publish its own telemetry should not automatically subscribe to another device's commands.
HTTP authorization usually maps identity to routes, resources, methods, and sometimes object ownership. A valid token should not grant access to arbitrary device identifiers supplied in a URL. Gateways make rate limits and request logs familiar, but they do not replace per-device authorization.
Payload validation and versioning are application responsibilities in both cases. MQTT is payload-agnostic, and HTTP media types describe representation formats without proving that fields are safe or compatible.
Operating the backend is part of the choice
A broker centralizes connections, subscriptions, fan-out, retained messages, session state, and delivery queues. That is useful infrastructure and a concentrated operational responsibility. Teams need limits for connections, inflight messages, packet size, session duration, retained data, topic cardinality, slow consumers, and authorization checks. They also need metrics that separate client churn, authentication failures, quota rejection, downstream lag, and broker saturation.
An HTTP platform centralizes routes, certificates, authentication, rate limits, request bodies, status codes, and service dependencies. It needs bounded timeouts, body limits, idempotency behavior, load shedding, and observability across gateways and application services. Polling can create predictable load or a wasteful baseline depending on interval and fleet size.
Choose the failure domain you can understand. A protocol that saves firmware work while creating opaque backend state is not simpler. A familiar API that forces constant polling for urgent commands is not simpler either.
A hybrid architecture is often more honest
One device does not require one application protocol for every responsibility. A sensible split might use MQTT for live telemetry and asynchronous desired-state events, while HTTP handles firmware artifacts, large diagnostic uploads, provisioning exchanges, or explicit resource reads. Another product may use HTTP for all device communication and publish events internally after the API accepts them.
The hybrid choice has a cost: two clients, two authorization surfaces, two connection policies, and more integration tests. Use it when responsibilities are genuinely different, not because the team avoided making a decision. Keep identity, schema versions, event identifiers, and observability consistent across both paths.
Decision checklist
Before selecting or changing the protocol, write down:
- dominant interaction: event, command, resource operation, stream, or artifact transfer;
- direction and fan-out of each message class;
- connection and power duty cycle;
- offline queue owner, capacity, expiry, and overflow behavior;
- duplicate, ordering, and idempotency rules;
- meaning of protocol and application acknowledgements;
- latency required for server-originated actions;
- identity, authorization, and credential lifecycle;
- backend state, quotas, observability, and on-call ownership;
- measured behavior during reconnect and fleet-wide recovery.
Choose MQTT when brokered publish/subscribe, live bidirectional events, and decoupled fan-out are central and you can operate their state. Choose HTTP when bounded resource operations, sparse or batch transfer, and explicit request/response semantics make the system clearer. Use both only when the responsibilities remain easier to explain after the split.
The production test is not whether the device can publish or receive a 200 response. It is whether the complete system can explain what happened when connectivity disappears, a message is repeated, a command expires, or thousands of devices return at once.