Submitted:
29 June 2026
Posted:
30 June 2026
You are already at the latest version
Abstract
Keywords:
1. Introduction
- A formal model of an asynchronous, non-blocking edge API gateway pipeline, including closed-form expressions for throughput, queuing delay, and resource utilization under the M/M/c queuing model.
- A reference architecture implementing the formal model, with annotated design decisions for back-pressure propagation, lock-free request routing, and adaptive timeout management.
- Empirical benchmark results comparing the proposed framework against a synchronous baseline across five load levels, two upstream failure scenarios, and three hardware configurations.
- A reproducibility package including all benchmark scripts, configuration files, and raw result data.
2. Related Work
2.1. Synchronous vs. Asynchronous Gateway Architectures
2.2. Back-Pressure in Distributed Systems
2.3. Edge API Gateway Performance
2.4. Formal Queuing Models for Gateway Systems
3. Formal Model
3.1. System Model
- : Inbound connection acceptance and TLS termination
- : Request parsing and authentication
- : Route matching and load balancer selection
- : Upstream connection pool management and proxying
- : Response transformation and client write
3.2. Throughput Bound
3.3. M/M/c Queuing Model
3.4. Back-Pressure Model
3.5. Resource Utilization
4. Reference Architecture
4.1. Overview

4.2. Event Loop Design
- I/O polling: Call epoll_wait (or io_uring_submit/io_uring_wait) with a 1 ms timeout to collect pending I/O readiness events.
- Handler dispatch: For each ready event, dispatch the associated continuation (a non-blocking callback or coroutine resume) inline.
- Timer processing: Process expired timers (adaptive timeouts, back-pressure tick, health-check callbacks) in FIFO order.
4.3. Lock-Free Request Routing
4.4. Adaptive Timeout Management
5. Empirical Evaluation
5.1. Experimental Setup
- Load levels: 20%, 40%, 60%, 80%, 95% of synchronous baseline peak throughput
- Upstream failure scenarios: (a) single upstream timeout storm (50% of requests timeout at 100 ms), (b) complete upstream pool exhaustion
- Hardware configurations: 2-core, 4-core, 8-core (via CPU affinity pinning)
5.2. Throughput and Latency Results
| Metric | Synchronous baseline | Async framework | Improvement |
|---|---|---|---|
| Peak throughput (req/s) | 29,600 | 112,400 | 3.8× |
| P50 latency (ms) | 14.2 | 4.1 | 3.5× |
| P99 latency (ms) | 69.4 | 18.3 | 3.8× |
| P99.9 latency (ms) | 312 | 41.7 | 7.5× |
| CPU utilization (%) | 91.4 | 63.2 | −31% |
| Memory (RSS, MB) | 4,218 | 298 | 14.1× less |
| Error rate (%) | 2.3 | 0.01 | 230× |
| Cores | Peak throughput (req/s) | Theoretical (Eq. 2) | Efficiency |
|---|---|---|---|
| 2 | 54,800 | 56,200 | 97.5% |
| 4 | 112,400 | 112,400 | 100% |
| 8 | 198,600 | 224,800 | 88.3% |
5.3. Back-Pressure Validation
| Pool utilization | Median (ms) | P99 (ms) | Client error rate (%) |
|---|---|---|---|
| 0.70 | — | — | 0.00 |
| 0.80 | — | — | 0.00 |
| 0.85 (threshold) | 4.2 | 9.1 | 0.02 |
| 0.90 | 4.4 | 9.8 | 0.08 |
| 0.95 | 4.7 | 11.2 | 0.31 |
| 1.00 (full) | 5.1 | 13.4 | 1.40 |
5.4. Memory Utilization
| Active connections | Synchronous RSS (MB) | Async RSS (MB) | Theoretical async (Eq. 12, MB) |
|---|---|---|---|
| 100 | 358 | 261 | 260.4 |
| 1,000 | 1,268 | 264 | 260.0 |
| 5,000 | 5,416 | 276 | 276.0 |
| 10,000 | 10,512 | 296 | 300.0 |
| 20,000 | OOM | 336 | 340.0 |
5.5. Upstream Failure Resilience
6. Discussion
6.1. Validity of the M/M/c Model
6.2. NUMA and Heterogeneous Core Effects
6.3. Security Considerations
6.4. Limitations
7. Conclusions
Author Contributions
Funding
Data Availability Statement
Conflicts of Interest
References
- Shi, W.; Cao, J.; Zhang, Q.; Li, Y.; Xu, L. Edge computing: Vision and challenges. IEEE Internet Things Journal. 2016, 3(5), 637–646. [Google Scholar] [CrossRef]
- Li, H.; Ota, K.; Dong, M. Learning IoT in edge: Deep learning for the Internet of Things with edge computing. IEEE Netw. 2018, 32(1), 96–101. [Google Scholar] [CrossRef]
- Poutanen, T.; Helin, H.; Kutvonen, L. Comparing synchronous and asynchronous web service invocations. In Proceedings of the IEEE International Conference on Web Services (ICWS); IEEE: Washington DC, 2008; pp. 249–256. [Google Scholar]
- Reese, W. Nginx: The high-performance web server and reverse proxy. Linux J. 2008, 2008(173), 2. [Google Scholar]
- Fielding, R.T.; Taylor, R.N. Principled design of the modern web architecture. ACM Trans. Internet Technol. 2002, 2(2), 115–150. [Google Scholar] [CrossRef]
- Erb, B. Concurrent programming for scalable web architectures. Diploma thesis, Ulm University, 2012. [Google Scholar]
- Schmidt, D.C.; Stal, M.; Rohnert, H.; Buschmann, F. Pattern-Oriented Software Architecture. In Patterns for Concurrent and Networked Objects; Wiley: Chichester, 2000; Volume 2. [Google Scholar]
- Boner, J. Thinking about concurrency: Reactive programming and the Actor model. Proceedings of QCon New York. InfoQ, 2014. [Google Scholar]
- Grigorik, I. High Performance Browser Networking; O’Reilly Media: Sebastopol, CA, 2013. [Google Scholar]
- Reactive Streams Working Group. Reactive Streams Specification for the JVM, v1.0.4. 2022. Available online: https://www.reactive-streams.org/.
- Akka Team. Akka Streams documentation: Back-pressure explained. Lightbend. 2023. Available online: https://doc.akka.io/docs/akka/current/stream/stream-introduction.html.
- Nguyen, D.T.; Kim, Y.; Tran, N.H. Performance evaluation of open-source API gateways for microservices architectures. IEEE Access. 2021, 9, 112345–112358. [Google Scholar]
- Kleinrock, L. Queuing Systems. In Theory; Wiley-Interscience: New York, 1975; Volume 1. [Google Scholar]
- Menascé, D.A.; Almeida, V.A.F. Capacity Planning for Web Services: Metrics, Models, and Methods; Prentice Hall: Upper Saddle River, NJ, 2001. [Google Scholar]
- Cloudwego. netpoll: A high-performance non-blocking I/O networking framework. GitHub/cloudwego. 2023. Available online: https://github.com/cloudwego/netpoll.
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).