Submitted:
22 August 2026
Posted:
24 August 2026
You are already at the latest version
Abstract
Distributed manufacturing requires communication that remains observable and recoverable as device populations, publish rates, and site boundaries grow. This paper presents a Message Queuing Telemetry Transport (MQTT) reference architecture and evaluates its communication core in a controlled, reproducible single-host experiment. The architecture combines a governed multi-factory topic namespace, message-class-specific Quality of Service (QoS), bounded queues, idempotent command handling, and randomized exponential reconnect control. Eclipse Mosquitto 2.1.2 was exercised with MQTT 5 clients on an Apple M3 Pro host. Forty-two steady-state trials covered 10–1000 concurrent publishers, aggregate loads of 25–2500 messages/s, and QoS 0–2 using 512-byte JSON payloads; six additional trials imposed a nominal 5 s broker outage on 250 publishers. Across 406,350 steady-state and 39,000 outage-test messages, measured delivery was 100%, with no duplicates or publisher errors. At 1000 publishers and 1000 messages/s, trial-level mean latency was 0.63 ± 0.61 ms and P95 latency was 1.43 ± 1.22 ms (mean ± 95% confidence interval), while broker CPU averaged 2.79 ± 0.64%. At 2500 messages/s, mean broker CPU was 4.81 ± 0.60%. QoS 2 increased mean broker CPU from 1.72 ± 0.10% at QoS 0 to 3.39 ± 0.20% under the matched 500-message/s workload. During recovery, jittered exponential backoff reduced peak connection attempts from 237.3 ± 54.5 to 57.0 ± 6.6 per 100 ms, although P95 reconnection increased from 0.13 ± 0.12 s to 0.93 ± 0.04 s. These results validate functional scaling and reconnect shaping within the tested loopback environment, not hard real-time or factory-wide performance.
Keywords:
MQTT
; industrial internet of things
; distributed manufacturing
; smart manufacturing
; edge computing
; Industry 4.0
1. Introduction
Industry 4.0 connects physical production assets with computation, data services, and operational decision systems. In distributed manufacturing, the relevant assets may span production cells, buildings, suppliers, and geographically separated factories. This distribution changes communication from a local integration problem into a lifecycle problem involving identity, routing, observability, failure recovery, and governance. Surveys of Industry 4.0 and cyber-physical production systems identify interoperable connectivity and timely data exchange as prerequisites for decentralized control and analytics [1,2,3]. Industrial Internet of Things (IIoT) research further emphasizes that constrained endpoints must coexist with legacy equipment, heterogeneous networks, and stronger availability and security requirements than are typical in consumer deployments [4,5].
Message Queuing Telemetry Transport (MQTT) is attractive in this setting because it separates publishers from subscribers through a broker and represents information flows as topic-addressed messages. MQTT 5.0 adds reason codes, session expiry, message expiry, topic aliases, request/response metadata, and richer flow control, while MQTT 3.1.1 remains widely deployed [6,7]. The protocol’s compact control packets and asynchronous publish/subscribe interaction reduce endpoint complexity, but neither version supplies hard real-time guarantees, authorization policy, durable application semantics, or proof that a machine executed a command. Those properties arise from the surrounding architecture.
Scaling therefore cannot be reduced to the number of TCP connections. Broker bottlenecks emerge from authentication, subscription matching, fan-out, persistence, retransmission, retained-message handling, encryption, and application consumers. Broad wildcard subscriptions can amplify routing work; QoS 1 and 2 introduce state and additional exchanges; disconnected persistent sessions can accumulate backlogs; and an outage can synchronize reconnect attempts into a storm. Published work confirms that MQTT cluster throughput can scale sub-linearly and that broker bridging, distributed placement, and edge/cloud partitioning materially affect latency and resource demand [8,9,10,11,12,13].
Prior work has benchmarked individual brokers, proposed distributed broker designs, or described IIoT architectures at a high level. Fewer studies connect message taxonomy, namespace governance, duplicate-safe control, bounded offline buffering, and reconnect shaping with a reproducible manufacturing workload. This study addresses that gap by coupling an explicit architecture with a measured baseline whose raw records, configuration, and analysis workflow are retained.
The objectives are to: (1) define a scalable MQTT communication architecture for multi-site manufacturing; (2) quantify latency, throughput, delivery, bandwidth, and broker resources as connection count and publish rate increase; (3) compare QoS 0–2 under a matched workload; and (4) quantify the restoration-speed and connection-burst tradeoff between fixed retry and jittered exponential backoff. The research questions are:
RQ1: How do latency, throughput, delivery, and broker resource use change from 10 to 1000 concurrent publishers at one message/s per publisher?
RQ2: How does aggregate publish rate affect latency and resource use for 250 concurrent publishers?
RQ3: What latency, bandwidth, and broker-resource differences arise among QoS 0, 1, and 2 under a matched 500-message/s workload?
RQ4: During a shared broker outage, how does jittered exponential backoff alter reconnection concentration and recovery time relative to fixed 100 ms retry?
RQ5: What do the measured results imply for the architecture, and which conclusions remain untested outside the single-host baseline?
The contributions are a governed multi-factory topic and message-policy model; an executable MQTT 5 benchmark spanning client scale, rate, QoS, and outage recovery; a message-level dataset covering 445,350 generated messages; and evidence that reconnect jitter can substantially reduce synchronized connection demand while preserving queued-message delivery in the tested setting.
2. Related Work
2.1. MQTT Semantics and Performance
MQTT delivers messages through broker-maintained subscriptions. QoS 0 provides best-effort transfer without protocol acknowledgment; QoS 1 retransmits until acknowledged and may duplicate a message; QoS 2 uses a four-step handshake to prevent duplicate delivery at the MQTT receiver [6,7]. These semantics apply to message transfer, not to physical actuation. A survey by Mishra and Kertesz shows that performance depends on broker, library, persistence, payload, network, and workload, making cross-study comparisons difficult when configurations are incompletely reported [8]. Detti et al. found sub-linear scaling in clustered MQTT publish/subscribe workloads, demonstrating that node count alone is not a capacity model [9]. Longo et al. designed an advanced broker for distributed scenarios [10], while distributed placement and ring-topology work examined how broker location and inter-broker routing change the path and load [11,12].
More recent edge/cloud and high-availability designs seek to reduce central memory pressure or tolerate broker failure [13,14,15]. These contributions support distribution, but manufacturing deployments still need explicit rules for which data may be dropped, replayed, expired, or compacted. A throughput improvement is unsafe if stale commands survive an outage, and perfect transport delivery is wasteful for telemetry that will be superseded within one sampling interval.
2.2. Industrial IoT, Edge Computing, and Manufacturing State
Reference IIoT literature describes an environment with heterogeneous field devices, gateways, enterprise services, and cyber-physical feedback [4,5,16]. Edge computing places computation and storage near data sources to limit wide-area dependency and latency [17,18]. In manufacturing, this permits local protocol adaptation, validation, buffering, and safety interlocks while central services provide fleet analytics and cross-site coordination. The design boundary is important: the cloud may issue a desired state, but only a local controller with current process context should authorize hazardous motion.
Digital-twin research formalizes synchronization between a physical asset and its digital representation [19,20,21,22,23]. For MQTT, a pragmatic device twin can be narrower than a full simulation model: desired state, reported state, configuration revision, firmware version, health, connectivity, and last-seen time. The useful contribution is not the label ‘twin’ but deterministic convergence rules, ownership of each field, conflict resolution, and traceable transitions.
2.3. Security and Interoperability
MQTT does not itself encrypt traffic or define a complete identity lifecycle. Transport Layer Security (TLS) 1.3 provides confidentiality and integrity in transit [24]; device-specific authentication and topic-level authorization must be enforced by broker and provisioning infrastructure. Zero-trust guidance recommends explicit, continuously evaluated access decisions rather than trust derived from network location [25]. Industrial control guidance additionally separates zones and conduits and requires component-level security capabilities [26]. When plant interoperability requires information models or deterministic control, MQTT may complement rather than replace OPC Unified Architecture (OPC UA), whose service and information models are standardized in IEC 62541 [27].
General IoT surveys and protocol comparisons show why no single protocol or topology meets every constraint [28,29,30]. The research gap is consequently architectural and empirical: a scalable MQTT design must bind protocol semantics to manufacturing message classes, control replay and expiry, preserve local autonomy, and expose a load/recovery protocol that another laboratory can reproduce.
3. Materials and Methods
3.1. System Requirements
The architecture shall transport near-real-time telemetry, reported machine state, alerts, configuration, and remote commands; identify every client and site; enforce least-privilege topic access; preserve required data during bounded outages; monitor connection and application health; and persist selected data for audit and analytics. Non-functional requirements are scalable connection and routing capacity, bounded tail latency, controlled bandwidth, fault isolation, recoverability, secure provisioning, and observable degradation. Acceptance limits are deployment-specific and must be defined by message consequence: telemetry may tolerate controlled sampling loss, whereas safety-relevant state and command workflows require explicit delivery, expiry, and execution criteria.
3.2. Proposed System Architecture
Figure 1 shows four layers. The device layer contains machines, programmable logic controllers (PLCs), sensors, actuators, and embedded MQTT clients. A device publishes measurements and reported state; command execution remains inside the controller’s validated control logic. The edge/gateway layer adapts legacy protocols, validates schemas, timestamps messages, performs permitted local rules, and stores a bounded disk-backed queue. Queue admission is message-class aware: critical state and events are preserved preferentially, while stale periodic telemetry may be compacted or expired.
The communication layer consists of one broker or a broker pool, a device registry, certificate and credential services, access-control lists (ACLs), retained state, session storage, and monitoring. Load balancing must preserve the selected broker’s session and subscription model; an arbitrary TCP load balancer does not create a coherent cluster. The cloud/application layer contains ingestion services, a time-series database, twin/state service, alerting, dashboards, analytics, and a command service. Commands include a unique identifier, creation and expiry times, target, type, parameters, configuration revision, and correlation identifier. The device publishes receipt and terminal execution outcomes separately.
3.3. MQTT Topic Architecture
The canonical namespace is factory/{factory_id}/area/{area_id}/machine/{machine_id}/{message_class}. Stable identifiers are used instead of display names. The extra area level permits operational partitioning without encoding network addresses. Publishers are denied write access outside their assigned subtree. Application subscriptions use the narrowest feasible filter; a global factory/+/.../# subscription is reserved for audited administration because it increases fan-out and authorization scope. Table 1 assigns defaults that must be validated against the chosen broker and workload.
3.4. Quality of Service Strategy
QoS is selected by consequence and freshness. High-rate, replaceable telemetry uses QoS 0 when occasional loss is acceptable; QoS 1 is used when gaps impair analysis. Machine state, availability, configuration revisions, and most alerts use QoS 1. A critical command may use QoS 1 with an idempotency key and explicit application acknowledgment. QoS 2 is considered only when its extra handshake and broker/client state are justified by measured risk and capacity. It prevents duplicate MQTT delivery to the receiver under the protocol assumptions, but it does not guarantee that an actuator executed once. Devices therefore store recently processed command identifiers and return cached outcomes for duplicates.
3.5. Device State and Twin Management
Each cloud-side representation contains desired and reported state, schema version, configuration revision, firmware version, connectivity, health, last communication timestamp, and operating status. Ownership is field-specific: applications write desired state; devices write reported state. On reconnect, the device first publishes availability, then reported configuration and operating state with a monotonically increasing device boot/session identifier. The twin service compares revisions, rejects stale writes, and either republishes the current desired revision or records convergence. Timestamps use Coordinated Universal Time, but ordering does not rely on wall-clock time alone because clocks may step.
3.6. Reliability Mechanisms
Clients use keep-alive and a retained Last Will and Testament (LWT) availability message. Clean-start and session-expiry settings are explicit; persistence is enabled only for subscriptions that require offline delivery. The edge queue is bounded by bytes, age, and message count, with a dead-letter record for rejected critical events. Every message includes message_id, source_id, boot_id, sequence, sent_at, schema_version, and optional expires_at. Consumers deduplicate by source and identifier within a documented window.
Reconnect delay after attempt k is d_k = min(d_max, d_0 2^k) + U(0,J_k), where U is a uniform jitter and J_k is the configured jitter span. The client resets k only after a stable connection interval, preventing rapid oscillation. Backlog replay is rate limited and interleaved with live state so recovery traffic cannot starve current operations. Retained messages carry current compact state, not event history. Broker redundancy is valid only when session, subscription, retained, and inflight state behavior has been verified for the implementation.
3.7. Security Architecture
All non-isolated links use TLS with server authentication and, where operationally feasible, mutual certificate authentication. Each device receives a unique identity during secure provisioning; shared fleet passwords are prohibited. Broker ACLs restrict publish and subscribe actions by canonical topic subtree. Private keys are stored in protected hardware when available, credentials rotate with overlap, revoked devices are denied promptly, and administrative APIs are separated from data-plane listeners. Payload validation, maximum packet size, connection quotas, inflight limits, audit logs, and anomaly detection limit abuse. These controls are infrastructure capabilities around MQTT, not guarantees of the MQTT protocol itself [24,25,26].
4. Experimental Setup
4.1. Test Environment and Factors
Experiments ran on a MacBook Pro with an Apple M3 Pro processor (11 CPU cores: five performance and six efficiency), 18 GB unified memory, and macOS 26.6.1. Eclipse Mosquitto 2.1.2_1 listened on 127.0.0.1:18884 with MQTT 5, anonymous local access, persistence disabled, a 1000-message inflight limit, a 20,000-message queued limit, and a 1 MB maximum message size. Steady-state publishers used aiomqtt 2.5.1; the wildcard measurement subscriber and recovery observer used Eclipse Paho MQTT 2.1.0. Every application payload was a deterministic 512-byte UTF-8 JSON record containing trial, client, sequence, and monotonic send timestamp. TLS, authentication, disk persistence, retained messages, and wide-area impairment were deliberately excluded to establish a communication-core baseline.
The scalability experiment used 10, 50, 100, 250, 500, and 1000 concurrent publishers at 1 message/s per client and QoS 1 for 20 s. The frequency experiment used 250 publishers at 0.1, 0.2, 1, 2, and 10 messages/s per client (25–2500 messages/s aggregate), QoS 1, for 20 s at the two lowest rates and 30 s otherwise. The QoS comparison used 250 publishers at 2 messages/s each for 30 s at QoS 0, 1, and 2. Each cell was repeated three times. A single wildcard subscriber received every message. Broker process CPU and resident set size and loopback-interface byte counters were sampled every 250 ms.
4.2. Metrics
For message i, one-way application latency was L_i = t_receive,i − t_send,i (1). Both timestamps used time.perf_counter_ns() on the same host, avoiding cross-host clock error. Trial-level mean, median, P95, and P99 were calculated from per-message observations. This latency includes client, loopback TCP/IP, broker, subscription matching, and subscriber callback time, but not a physical or wide-area network.
Delivery rate was R_d = 100N_unique_received/N_published (2); duplicate rate was R_dup = 100(N_received − N_unique_received)/N_unique_received (3); and throughput was T = N_unique_received/Δt (4). Loopback traffic was estimated from interface-byte deltas and therefore includes both directions and protocol overhead. Broker CPU was measured as process CPU percentage and memory as peak resident set size. Recovery trials additionally recorded actual outage, unique delivery, duplicates, time from broker restoration to reconnect, and the maximum connection attempts observed in any 100 ms bin during the first 2 s after restoration.
4.3. Trial Procedure and Statistical Analysis
Each steady-state trial started a fresh broker, connected and subscribed the observer, connected all publishers, and waited 1 s before measurement. Publishers used scheduled monotonic release times; a 3 s drain period followed transmission. Treatments were run in ascending factor order and repeated three times. Summary values are means of the three independent trial statistics with two-sided 95% Student-t confidence intervals. Because n = 3 yields imprecise interval estimates, individual raw message records are retained and interpretation emphasizes effect magnitude and consistency rather than null-hypothesis testing. Instrumentation-invalid recovery attempts, identified by observer readiness failures, were archived separately and excluded before final analysis.
5. Evaluation Scenarios
5.1. Scenario 1—Scalability Under Increasing Device Count
Concurrent publishers increased from 10 to 1000 at one QoS 1 message/s per publisher. The experiment measured delivered throughput, latency distribution, delivery, duplicates, broker CPU and memory, and loopback traffic.
5.2. Scenario 2—Impact of Message Frequency
With 250 publishers fixed, per-client rate increased from 0.1 to 10 messages/s. This isolated message-processing demand from connection count across aggregate offered loads of 25–2500 messages/s.
5.3. Scenario 3—MQTT QoS Comparison
Identical 512-byte messages were published at QoS 0, 1, and 2 with 250 publishers and an aggregate offered load of 500 messages/s. Broker state was cleared between trials.
5.4. Scenario 4—Network Failure and Recovery
The broker was stopped for a nominal 5 s while 250 clients continued generating one message/s into bounded 100-message local queues. On restoration, each client reconnected and drained its queue. The observer established subscription readiness before backlog replay so measurement infrastructure did not misclassify replayed records as loss.
5.5. Scenario 5—Simultaneous Device Reconnection
Recovery compared fixed 100 ms retry with randomized exponential backoff: min(2.0 s, 0.1×2^k) plus uniform jitter up to 0.25 s. Each policy was repeated three times. Reconnection timing was measured before the observer-readiness gate; therefore the gate protected delivery measurement without improving the reported reconnect metric.
6. Results
6.1. Scalability with Concurrent Publishers
All 18 scalability trials achieved 100% unique delivery without duplicates or publisher errors. Delivered throughput equaled offered load at every treatment. At 1000 clients, throughput was 1000 messages/s, trial-level mean latency was 0.63 ± 0.61 ms, P95 was 1.43 ± 1.22 ms, broker CPU was 2.79 ± 0.64%, peak resident memory was 9.38 ± 0.04 MiB, and measured loopback traffic was 21.40 ± 0.12 Mbit/s. The wide P99 interval at 1000 clients reflects one trial with a longer tail; P95 and mean latency remained low in all repetitions.
Table 2.
Scalability results (mean ± 95% CI across three trials).
| Clients | Msg/s | Mean latency (ms) | P95 (ms) | CPU (%) | Peak RSS (MiB) |
|---|---|---|---|---|---|
| 10 | 10 ± 0 | 0.93 ± 0.39 | 1.75 ± 0.97 | 0.08 ± 0.04 | 4.54 ± 0.02 |
| 50 | 50 ± 0 | 0.98 ± 0.26 | 1.96 ± 0.66 | 0.36 ± 0.08 | 4.71 ± 0.09 |
| 100 | 100 ± 0 | 0.82 ± 0.24 | 1.82 ± 0.76 | 0.60 ± 0.10 | 4.97 ± 0.07 |
| 250 | 250 ± 0 | 0.64 ± 0.31 | 1.39 ± 0.63 | 1.11 ± 0.38 | 5.68 ± 0.04 |
| 500 | 500 ± 0 | 0.59 ± 0.37 | 1.32 ± 0.82 | 1.72 ± 0.40 | 6.94 ± 0.04 |
| 1000 | 1000 ± 0 | 0.63 ± 0.61 | 1.43 ± 1.22 | 2.79 ± 0.64 | 9.38 ± 0.04 |
Figure 2.
Application latency versus concurrent QoS 1 publishers. Error bars show 95% confidence intervals across three trial-level statistics.
Figure 2.
Application latency versus concurrent QoS 1 publishers. Error bars show 95% confidence intervals across three trial-level statistics.

Figure 3.
Delivered throughput versus offered load in the scalability experiment.

6.2. Publish-Frequency Effects
The frequency experiment also produced 100% measured delivery without duplicates. With 250 connected clients, delivered throughput tracked offered load from 25 to 2500 messages/s. At 2500 messages/s, mean latency was 0.53 ± 0.18 ms, P95 latency was 1.12 ± 0.39 ms, broker CPU was 4.81 ± 0.60%, and loopback traffic was 51.71 ± 0.19 Mbit/s. Within this tested range, increasing load did not produce a latency knee; broker CPU and interface traffic increased while latency remained below the lower-rate trial means.
Table 3.
Publish-frequency results with 250 QoS 1 publishers (mean ± 95% CI).
| Offered msg/s | Mean latency (ms) | P95 (ms) | CPU (%) | Loopback (Mbit/s) |
|---|---|---|---|---|
| 25 | 1.05 ± 0.06 | 1.92 ± 0.38 | 0.25 ± 0.03 | 0.65 ± 0.00 |
| 50 | 0.95 ± 0.14 | 1.71 ± 0.47 | 0.38 ± 0.09 | 1.20 ± 0.01 |
| 250 | 0.92 ± 0.32 | 2.00 ± 0.89 | 1.24 ± 0.27 | 5.35 ± 0.02 |
| 500 | 1.01 ± 0.87 | 2.65 ± 2.49 | 1.73 ± 1.15 | 10.45 ± 0.34 |
| 2500 | 0.53 ± 0.18 | 1.12 ± 0.39 | 4.81 ± 0.60 | 51.71 ± 0.19 |
Figure 4.
Broker CPU and loopback traffic versus aggregate publish rate. Error bars show 95% confidence intervals.
Figure 4.
Broker CPU and loopback traffic versus aggregate publish rate. Error bars show 95% confidence intervals.

6.3. QoS Comparison
All QoS treatments delivered 500 unique messages/s with no measured loss or duplicates. P95 latency rose from 1.92 ± 0.65 ms at QoS 0 to 2.23 ± 0.13 ms at QoS 2. Mean broker CPU increased from 1.72 ± 0.10% to 3.39 ± 0.20%, while loopback traffic increased from 9.06 ± 0.02 to 13.44 ± 0.04 Mbit/s. Thus QoS 2 imposed a clear resource and network cost without a delivery advantage in this loss-free local setting.
Table 4.
QoS comparison at 500 messages/s (mean ± 95% CI).
| QoS | Mean latency (ms) | P95 (ms) | CPU (%) | Loopback (Mbit/s) |
|---|---|---|---|---|
| 0 | 0.79 ± 0.11 | 1.92 ± 0.65 | 1.72 ± 0.10 | 9.06 ± 0.02 |
| 1 | 0.88 ± 0.11 | 2.09 ± 0.47 | 2.33 ± 0.21 | 10.53 ± 0.04 |
| 2 | 1.10 ± 0.08 | 2.23 ± 0.13 | 3.39 ± 0.20 | 13.44 ± 0.04 |
Figure 5.
P95 latency and broker CPU by QoS under a matched workload.

6.4. Broker-Outage Recovery
Across six recovery trials, all 39,000 generated messages were observed exactly once and every bounded client queue drained. The actual outage averaged 5.063 ± 0.008 s for fixed retry and 5.059 ± 0.005 s for jittered backoff. Fixed retry restored most clients rapidly but concentrated attempts: mean peak demand was 237.3 ± 54.5 attempts/100 ms. Jitter reduced this peak by approximately 76% to 57.0 ± 6.6 attempts/100 ms, while mean P95 reconnect time increased from 0.13 ± 0.12 s to 0.93 ± 0.04 s.
Table 5.
Broker-outage recovery results (mean ± 95% CI across three trials).
| Retry policy | Outage (s) | Delivered | Median reconnect (s) | P95 reconnect (s) | Peak attempts/100 ms |
|---|---|---|---|---|---|
| Fixed 100 ms | 5.063 ± 0.008 | 19500/19500 | 0.097 ± 0.003 | 0.126 ± 0.118 | 237.3 ± 54.5 |
| Jittered exponential | 5.059 ± 0.005 | 19500/19500 | 0.658 ± 0.018 | 0.928 ± 0.043 | 57.0 ± 6.6 |
Figure 6.
Recovery-speed and connection-burst tradeoff. Error bars show 95% confidence intervals.

7. Discussion
RQ1 is answered within the loopback baseline: Mosquitto maintained exact offered throughput and 100% observed delivery through 1000 publishers, with low mean and P95 latency and modest process CPU. Memory increased with connected clients, from 4.54 ± 0.02 MiB at 10 clients to 9.38 ± 0.04 MiB at 1000. These values are capacity observations for one configuration, not a universal MQTT limit.
For RQ2, aggregate rate increased by two orders of magnitude without a latency knee. The dominant measured changes were broker CPU and loopback traffic. The lower mean latency at the highest rate likely reflects batching and continuously active event-loop paths; it should not be interpreted as evidence that load generally improves latency. A distributed network with TLS, persistence, slow subscribers, or storage contention may behave differently.
RQ3 shows why QoS should follow consequence. In a loss-free local environment, QoS 0–2 had identical measured delivery, while QoS 2 approximately doubled broker CPU relative to QoS 0 and increased interface traffic by about 48%. QoS 1 remains the defensible default for state, alerts, configuration, and commands when consumers are duplicate-safe; QoS 2 requires a risk-based justification. MQTT acknowledgment still does not prove physical execution.
For RQ4, synchronized fixed retry created a concentrated arrival burst. Jittered exponential backoff reduced the mean peak attempt count by approximately 76% and produced consistent recovery across repetitions, at the cost of about 0.80 s greater mean P95 reconnect time. This is the intended systems tradeoff: a small restoration delay protects authentication, session restoration, and subscription processing from a thundering herd.
RQ5 is addressed by combining the measurements with architectural safeguards. The tested broker had substantial headroom, so these data do not justify clustering at 1000 local publishers or 2500 messages/s. Edge federation or clustering should be considered only after representative tests with security, persistence, real fan-out, and network conditions breach declared latency, delivery, recovery, or resource thresholds. Published evidence of sub-linear cluster scaling cautions against assuming that additional nodes automatically produce proportional capacity [9].
8. Limitations
The experiment used one broker implementation, one host, loopback networking, simulated clients, a single wildcard subscriber, one 512-byte JSON payload, three repetitions per treatment, and short measurement windows. It disabled TLS, authentication, persistence, retained messages, disk-backed broker queues, multiple subscribers, and network loss or delay. CPU and memory are process-level observations; loopback byte counters include both directions and may include small unrelated host traffic. The recovery observer gated backlog replay until resubscription to prevent measurement loss, a control that isolates queue/retry behavior but does not represent every production consumer. No physical machine, programmable controller, hazardous actuation, cloud link, multi-broker topology, or long-duration factory workload was tested. The data therefore validate the software communication core under controlled conditions, not industrial safety, security, hard real-time deadlines, or universal scalability.
9. Future Work
Future work should repeat the matrix across Linux hosts and multiple broker implementations; add TLS, mutual authentication, access-control checks, persistence, retained state, fan-out, larger payloads, and slow consumers; and introduce controlled delay, jitter, loss, bandwidth constraints, and multi-site links. Longer endurance runs should quantify storage growth and resource leakage. The recovery experiment should compare queue policies, session persistence, message expiry, backlog rate limiting, and multiple outage durations. Finally, physical gateways and PLC-adjacent systems should validate timing, semantic reconciliation, safety boundaries, and operational maintainability before factory deployment.
10. Conclusion
This study couples a governed MQTT architecture for distributed manufacturing with an executable baseline experiment. Across 42 steady-state trials, all 406,350 published messages were received uniquely; the tested broker sustained 1000 concurrent one-message/s publishers and 2500 messages/s from 250 publishers with low application latency and modest CPU use. QoS 2 increased CPU and traffic without improving delivery in the loss-free loopback setting. Across six outage trials, all 39,000 queued messages were recovered, and jittered exponential backoff reduced peak connection attempts by approximately 76% while increasing P95 reconnection by about 0.80 s. The results support message-class-specific QoS, bounded client queues, duplicate-safe processing, and randomized reconnect control. They do not establish hard real-time or factory-wide performance; representative security, network, persistence, fan-out, and physical-device testing remains necessary before deployment.
Author Contributions
Nata Sulakvelidze and Anzori Kuparadze jointly contributed to the conceptualization of the architecture, methodology development, technical review, and manuscript revision. Both authors reviewed and approved the final manuscript.
Funding
No specific external grant was received for this study.
Institutional Review Board Statement
The study used software-generated MQTT traffic and did not involve human participants, identifiable personal data, animals, or operational factory equipment; ethics approval was not applicable.
Data Availability Statement
Raw per-message latency records, trial summaries, broker configuration, benchmark source code, analysis source code, and processed aggregate tables are supplied with this study and are available from the corresponding author on reasonable request.
Acknowledgments
None.
Conflicts of Interest
The authors are officers of G3D LLC, as disclosed in their affiliations. The manuscript presents a technology-neutral reference architecture and does not evaluate or endorse a commercial product.
References
- Lu, Y. (2017). Industry 4.0: A survey on technologies, applications and open research issues. Journal of Industrial Information Integration, 6, 1–10. [CrossRef]
- Xu, L. D., Xu, E. L., & Li, L. (2018). Industry 4.0: State of the art and future trends. International Journal of Production Research, 56(8), 2941–2962. [CrossRef]
- Lee, J., Bagheri, B., & Kao, H.-A. (2015). A cyber-physical systems architecture for Industry 4.0-based manufacturing systems. Manufacturing Letters, 3, 18–23. [CrossRef]
- Boyes, H., Hallaq, B., Cunningham, J., & Watson, T. (2018). The industrial internet of things (IIoT): An analysis framework. Computers in Industry, 101, 1–12. [CrossRef]
- Sisinni, E., Saifullah, A., Han, S., Jennehag, U., & Gidlund, M. (2018). Industrial Internet of Things: Challenges, opportunities, and directions. IEEE Transactions on Industrial Informatics, 14(11), 4724–4734. [CrossRef]
- OASIS. (2019). MQTT Version 5.0. OASIS Standard. https://docs.oasis-open.org/mqtt/mqtt/v5.0/mqtt-v5.0.html.
- OASIS. (2014). MQTT Version 3.1.1. OASIS Standard. https://docs.oasis-open.org/mqtt/mqtt/v3.1.1/mqtt-v3.1.1.html.
- Mishra, B., & Kertesz, A. (2020). The use of MQTT in M2M and IoT systems: A survey. IEEE Access, 8, 201071–201086. [CrossRef]
- Detti, A., Tropea, G., Rossi, G., Melazzi, N. B., & Salsano, S. (2020). Sub-linear scalability of MQTT clusters in topic-based publish-subscribe applications. IEEE Transactions on Network and Service Management, 17(3), 1954–1968. [CrossRef]
- Longo, E., Redondi, A. E. C., Cesana, M., Arcia-Moret, A., & Manzoni, P. (2023). Design and implementation of an advanced MQTT broker for distributed pub/sub scenarios. Computer Networks, 224, 109601. [CrossRef]
- Fawwaz, D. Z., Chung, S.-H., & Nishiyama, H. (2022). Optimal distributed MQTT broker and services placement for SDN-edge based smart city architecture. Sensors, 22(9), 3431. [CrossRef]
- Ohno, S., Hasegawa, G., & Murata, M. (2021). Distributed MQTT broker architecture using ring topology and its prototype. IEICE Communications Express, 10(10), 787–793. [CrossRef]
- Tseng, Y.-H., Wang, C., Wei, Y.-T., & Chiang, Y.-T. (2025). Cloud-edge MQTT messaging for latency mitigation and broker memory footprint reduction. PeerJ Computer Science, 11, e2741. [CrossRef]
- Shvaika, D., Shvaika, A., & Artemchuk, V. (2025). MQTT broker architectural enhancements for high-performance P2P messaging: TBMQ scalability and reliability in distributed IoT systems. IoT, 6(3), 34. [CrossRef]
- Chai, A., Yin, W., Lian, M., Sun, Y., Guo, C., Wang, L., & Fang, Z. (2025). DUA-MQTT: A distributed high-availability message communication model for the Industrial Internet of Things. Sensors, 25(16), 5071. [CrossRef]
- Al-Fuqaha, A., Guizani, M., Mohammadi, M., Aledhari, M., & Ayyash, M. (2015). Internet of Things: A survey on enabling technologies, protocols, and applications. IEEE Communications Surveys & Tutorials, 17(4), 2347–2376. [CrossRef]
- Shi, W., Cao, J., Zhang, Q., Li, Y., & Xu, L. (2016). Edge computing: Vision and challenges. IEEE Internet of Things Journal, 3(5), 637–646. [CrossRef]
- Satyanarayanan, M. (2017). The emergence of edge computing. Computer, 50(1), 30–39. [CrossRef]
- Tao, F., Zhang, H., Liu, A., & Nee, A. Y. C. (2019). Digital twin in industry: State-of-the-art. IEEE Transactions on Industrial Informatics, 15(4), 2405–2415. [CrossRef]
- Tao, F., Anwer, N., Liu, A., Wang, L., Nee, A. Y. C., Li, L., & Zhang, M. (2021). Digital twin towards smart manufacturing and Industry 4.0. Journal of Manufacturing Systems, 58, 1–2. [CrossRef]
- Li, L., Lei, B., & Mao, C. (2022). Digital twin in smart manufacturing. Journal of Industrial Information Integration, 26, 100289. [CrossRef]
- Qi, Q., & Tao, F. (2018). Digital twin and big data towards smart manufacturing and Industry 4.0: 360 degree comparison. IEEE Access, 6, 3585–3593. [CrossRef]
- Qi, Q., Tao, F., Zuo, Y., & Zhao, D. (2018). Digital twin service towards smart manufacturing. Procedia CIRP, 72, 237–242. [CrossRef]
- Rescorla, E. (2018). The Transport Layer Security (TLS) Protocol Version 1.3 (RFC 8446). Internet Engineering Task Force. [CrossRef]
- Rose, S., Borchert, O., Mitchell, S., & Connelly, S. (2020). Zero Trust Architecture (NIST SP 800-207). National Institute of Standards and Technology. [CrossRef]
- IEC. (2019). IEC 62443-4-2:2019, Security for industrial automation and control systems—Part 4-2: Technical security requirements for IACS components. International Electrotechnical Commission.
- IEC. (2020). IEC 62541-1:2020, OPC Unified Architecture—Part 1: Overview and concepts. International Electrotechnical Commission.
- Gubbi, J., Buyya, R., Marusic, S., & Palaniswami, M. (2013). Internet of Things (IoT): A vision, architectural elements, and future directions. Future Generation Computer Systems, 29(7), 1645–1660. [CrossRef]
- Naik, N. (2017). Choice of effective messaging protocols for IoT systems: MQTT, CoAP, AMQP and HTTP. In 2017 IEEE International Systems Engineering Symposium (ISSE), 1–7. [CrossRef]
- Wytrębowicz, J., Cabaj, K., & Krawiec, J. (2021). Messaging protocols for IoT systems—A pragmatic comparison. Sensors, 21(20), 6904. [CrossRef]
Figure 1.
Proposed MQTT-based distributed manufacturing architecture. Bidirectional arrows denote protocol flows, not authorization for direct cloud control of safety-critical actuation.
Figure 1.
Proposed MQTT-based distributed manufacturing architecture. Bidirectional arrows denote protocol flows, not authorization for direct cloud control of safety-critical actuation.

Table 1.
Proposed MQTT topic structure and message functions.
| Topic suffix | Publisher | Subscriber | Purpose | QoS | Retained | Typical frequency |
|---|---|---|---|---|---|---|
| telemetry | Device/gateway | Ingestion, analytics | Time-series samples | 0 or 1 | No | 100 ms–10 s |
| status | Device | Twin, dashboard | Operating/connectivity state | 1 | Yes | On change + heartbeat |
| alerts | Device/edge | Alert service | Diagnostic or process event | 1 | No* | Event driven |
| command | Command service | Target device | Desired action with expiry | 1; 2 only if justified | No | Event driven |
| command/ack | Device | Command service | Receipt/execution outcome | 1 | No | Per command |
| configuration/desired | Twin/config service | Target device | Versioned desired configuration | 1 | Yes | On change |
| configuration/reported | Device | Twin/config service | Applied revision and result | 1 | Yes | On change |
| health | Device/gateway | Monitoring | Resource and diagnostic health | 0 or 1 | Optional | 5–60 s |
| availability | Device LWT/device | Monitoring/twin | Unexpected offline / online | 1 | Yes | Connection change |
*Alerts are normally event records, not retained truth. A separate retained active-alarm summary may be used if clearing semantics are explicit. Topic aliases in MQTT 5.0 may reduce repeated-topic overhead on constrained links without changing authorization [6].
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.