Preprint
Article

This version is not peer-reviewed.

Link-Time Bytecode Quickening for Java Card Virtual Machines without Method-Component Expansion: A Formal and Analytical Study

A peer-reviewed version of this preprint was published in:
Computers 2026, 15(10), 643. https://doi.org/10.3390/computers15100643

Submitted:

31 August 2026

Posted:

01 September 2026

You are already at the latest version

Abstract
Java Card virtual machines operate in resource-constrained secure elements where repeated reference indirection can add latency to frequently executed instructions. This paper proposes a link-time bytecode quickening scheme for statically resolvable method-invocation and static-field instructions. After ordinary installation-time validation and symbolic resolution, the linker rewrites each selected three-byte instruction into an implementation-private direct-address form of the same length. The low three bits of an aligned opcode family encode the three most significant address bits, while the existing two operand bytes encode the remaining sixteen bits, yielding a 19-bit logical address space of 512 KiB. The study is formal and analytical rather than experimental. It defines the transformation and decoder, proves address round-trip correctness, establishes preservation of Method-component length and control-flow offsets, and gives a conditional semantic-equivalence argument under explicit linker-correctness, address-stability, decoder-correctness, address-range, and opcode-separation assumptions. A parametric execution-cost model shows that the gain equals the removed indirect address-formation work, adjusted for any decoder-cost difference. The proposal targets JCVM implementations that retain indirect post-link reference representations; dynamic dispatch instructions are excluded. The scheme preserves ordinary CAP-file input and application source code while requiring a modified linker and interpreter for the installed representation.
Keywords: 
;  ;  ;  ;  ;  ;  

1. Introduction

Java Card is a compact execution platform for smart cards, secure elements, payment cards, and UICC-based applications. Its virtual-machine and CAP-file architecture supports constrained memory, portability, controlled installation, and application isolation [1]. Java Card memory organization is implementation-specific, and published design work demonstrates that method bytecode, package metadata, and runtime structures can be arranged in different ways while preserving platform semantics [2]. In systems that execute from non-volatile memory, the number and structure of repeated dereferences can materially affect the constant cost of interpreted instructions.
Bytecode quickening is a well-established virtual-machine technique in which a general instruction is replaced by a specialized form after resolution or validation work has been completed. Prior JVM work used quickened bytecodes to avoid repeated class-initialization and constant-pool resolution barriers [3]. Formal reasoning about bytecode transformations is important because an optimization must preserve the behavior of the original program and the structural assumptions used by verification and execution [4]. Java Card research has also shown that implementation-specific opcode and memory representations are relevant to performance and security analysis [5].
Installation-time linking already resolves symbolic information needed to connect packages, classes, fields, and methods. This must be distinguished from the representation used after linking. A JCVM can complete symbolic validation during installation yet retain an indirect linked-reference format for execution. Such a format may require one or more table dereferences, address calculations, or a branch that distinguishes internal and external targets. Other implementations may already perform more aggressive quickening. This work therefore does not characterize all compliant JCVMs; it defines an optimization for implementations that still execute selected statically resolvable operations through an indirect post-link representation.
The central proposal is an implementation-private direct-address instruction produced by the linker after all ordinary checks have succeeded. The transformed opcode stores the high three address bits, and the existing two operand bytes store the low sixteen bits. The transformation is performed in place, so the instruction remains three bytes and no later instruction moves. The paper is intentionally theoretical: it does not depend on a particular EMV, transit, SIM Toolkit, or access-control applet and does not claim a universal measured speedup. The contribution consists of a precise transformation, explicit preconditions, formal correctness properties, and a parametric cost model.
The principal contributions are as follows:
  • A formal definition of link-time quickening for selected Java Card static-reference instructions.
  • A 19-bit direct-address encoding that preserves the original three-byte instruction length.
  • Proofs of address round-trip correctness, instruction-layout invariance, offset preservation, and conditional semantic equivalence.
  • A compatibility and security analysis separating CAP-file input compatibility from the modified installed representation.
  • A parametric execution-cost model that avoids unsupported assumptions about candidate applets or hardware latency.
The remainder of the paper defines the baseline and transformation, establishes formal properties, derives analytical cost bounds, and discusses compatibility, applicability, limitations, and future work.

2. Materials and Methods

2.1. Java Card Installation and Linking

A Java Card CAP file contains components describing executable methods, constant-pool entries, imports, exports, and reference locations that require processing during installation [1]. The installer and linker validate references, enforce package visibility and type constraints, and produce an executable installed representation. The proposed quickening step is inserted only after the ordinary linker has accepted a reference and computed its target under the platform memory map.
The transformation is not a replacement for specification-defined validation. The linker remains responsible for rejecting unresolved, inaccessible, malformed, out-of-range, or otherwise invalid references. Quickening changes only the representation used after successful resolution.

2.2. Baseline Execution Model

Let an indirect linked-reference implementation be a JCVM in which an installed statically resolvable instruction contains, or points to, a linked reference rather than the final executable or data address. During execution, the interpreter decodes the instruction and performs a bounded sequence of implementation-specific operations that may include a link-table or constant-pool-cache dereference, package or object-table access, internal/external path selection, and final target-address formation.
This definition deliberately avoids the stronger claim that every compliant JCVM traverses the original CAP-file Export component on every invocation. A platform that already replaces references with direct executable addresses would obtain little or no benefit. A platform that retains one or more non-volatile-memory indirections is the intended target.

2.3. Problem Statement

For an executed statically resolvable instruction i, let I_i denote the implementation-specific work required to convert its installed linked representation into the target address. The optimization objective is to remove I_i from the repeated execution path while preserving:
  • the result of the ordinary linker checks;
  • the observable Java Card semantics of the instruction;
  • the three-byte instruction length;
  • all subsequent Method-component byte offsets;
  • control-flow destinations and exception-handler ranges; and
  • the isolation and access-control decisions already enforced by the platform.

2.4. Supported and Excluded Instructions

The transformation applies only when the final target can be selected at link time. Table 1 lists the supported semantic groups. Short and reference static-field variants may share address-decoding logic because their direct-address formation is identical, although their value interpretation remains instruction-specific.
The instructions invokevirtual and invokeinterface are excluded because their final targets depend on runtime receiver type, class hierarchy, or interface dispatch. Any instruction whose target can change without re-linking is also outside the scope.

2.5. Preconditions

A source instruction may be quickened only if all of the following conditions hold:
  • The ordinary Java Card installation and linking checks have completed successfully.
  • The linker has determined a unique target address a for the installed lifetime of the instruction.
  • The address satisfies 0 <= a < 2^19.
  • The target addressing domain is stable: the target will not move unless the instruction is re-linked or restored to an indirect form.
  • The implementation reserves a non-conflicting private opcode family for the quickened instruction.
  • The interpreter recognizes both the ordinary installed representation and the quickened representation, allowing a safe fallback.

2.6. Three-Byte, 19-Bit Direct-Address Encoding

Each quickened semantic group is assigned an aligned base opcode B whose low three bits are zero. The eight byte values B through B + 7 form one private opcode family. For a resolved address a, the low three bits of the emitted opcode store bits 18-16 of a, and the two operand bytes store bits 15-0.
q = B OR ((a >> 16) AND 0x07)
b1 = (a >> 8) AND 0xFF, b2 = a AND 0xFF
The interpreter reconstructs the direct address as follows:
Decode(q, b1, b2) = ((q AND 0x07) << 16) OR (b1 << 8) OR b2
Because 2^19 = 524,288 bytes, the construction covers a 512 KiB logical address domain. The instruction remains one opcode byte plus two operand bytes.
Table 2. Proposed implementation-private opcode-family mapping.
Table 2. Proposed implementation-private opcode-family mapping.
Private base B Quickened semantic group Opcode family Address bits in opcode
0xC0 putstatic_b 0xC0-0xC7 a [18:16]
0xC8 putstatic_s / putstatic_a 0xC8-0xCF a [18:16]
0xD0 getstatic_b 0xD0-0xD7 a [18:16]
0xD8 getstatic_s / getstatic_a 0xD8-0xDF a [18:16]
0xE0 invokestatic 0xE0-0xE7 a [18:16]
0xE8 getstatic_i 0xE8-0xEF a [18:16]
0xF0 invokespecial 0xF0-0xF7 a [18:16]
0xF8 putstatic_i 0xF8-0xFF a [18:16]
The construction requires aligned, non-overlapping private opcode families. The specific hexadecimal allocation must be checked against the target Java Card edition and implementation before deployment.

2.7. Linker Transformation Algorithm

Let R be the set of reference locations selected for possible quickening. The ordinary linker already visits these locations during installation. The added transformation is linear in the number of selected locations.
Preprints 231084 i001
No byte is inserted or removed. If any precondition fails, the implementation retains its ordinary linked form. This makes the optimization incremental rather than mandatory.

2.8. Interpreter Execution

Preprints 231084 i002
The direct handler must preserve the original instruction-specific behavior, including operand-stack effects, value width, method-frame creation, return behavior, exception behavior, transaction semantics, and any platform checks intentionally performed at execution time. Only repeated address-formation work is removed.

3. Results

3.1. Definitions and Assumptions

Let P be an accepted installed program, r a statically resolvable reference in P, Resolve(P,r) the address returned by the ordinary compliant linker, and a = Resolve(P,r). Let Exec_std(P,r,s) denote execution of the platform’s ordinary linked representation from machine state s, and Exec_q(P,a,s) execution of the corresponding quickened instruction from the same state. The analysis uses the following assumptions:
  • A1 - Linker correctness: Resolve(P,r) identifies the method or static storage entity intended by r after visibility, type, and access checks.
  • A2 - Address stability: a remains the target address for the installed lifetime of the quickened instruction, or the platform re-links before movement.
  • A3 - Decoder correctness: the quickened interpreter applies the same instruction semantics after obtaining a.
  • A4 - Address-domain validity: 0 <= a < 2^19.
  • A5 - No opcode collision: the quickened opcode family is distinguishable from standard and other private instructions.

3.2. Address Round-Trip Correctness

Lemma 1.
For every address a satisfying 0 <= a < 2^19, Equations (1)-(3) reconstruct a exactly.
Proof of Lemma 1. Write a = h x 2^16 + l, where h = (a >> 16) AND 0x07 and 0 <= l < 2^16. The emitted opcode contains h in its low three bits because the low three bits of B are zero. The two operand bytes encode l in big-endian order. Equation (3) reconstructs h x 2^16 + l = a □

3.3. Instruction-Length and Offset Preservation

Lemma 2.
Replacing a supported three-byte source or installed instruction with the quickened form does not change the length of the Method component.
Proof of Lemma 2. The transformation overwrites exactly three existing bytes with one quickened opcode byte and two address bytes. It performs no insertion or deletion; therefore, the Method-component length is invariant □
Corollary 1.
All instruction start offsets after a transformed location remain unchanged. Consequently, branch displacements, exception-handler boundaries, and other structures whose meaning depends on Method-component byte positions remain valid, provided those structures were valid before transformation and do not interpret the overwritten operand as a symbolic reference after installation.

3.4. Conditional Semantic Preservation

Theorem 1.
For a supported instruction with reference r and resolved address a = Resolve(P,r), Exec_std(P,r,s) and Exec_q(P,a,s) produce the same observable Java Card state transition for every valid pre-state s under assumptions A1-A5.
Proof of Theorem 1. The ordinary execution path uses the accepted linked representation of r to obtain target a and then applies the instruction-specific semantics to a. By Lemma 1, the quickened instruction reconstructs the same a. By A3, both paths apply the same method-call or static-field operation after address formation. Linker-time access and type decisions are identical by A1, and a remains valid by A2. The difference is therefore confined to the removed address-formation mechanism and does not change the observable method, field, stack, heap, persistent-state, or exception result □
Table 3. Instruction-class semantic cases.
Table 3. Instruction-class semantic cases.
nstruction class Ordinary target formation Quickened action Preserved effect
invokestatic Linked reference to static method target Direct method address a Same method, arguments, frame, return value, and exceptions.
invokespecial Linked reference to fixed special target Direct method address a Same special target under the accepted class relation.
getstatic_* Linked reference to static storage address Direct storage address a Same value width and value read from the same storage entity.
putstatic_* Linked reference to static storage address Direct storage address a Same value width and value written to the same storage entity.

3.5. Analytical Execution-Cost Model

The transformation changes constant factors rather than asymptotic complexity: both a bounded indirect lookup and a direct lookup are O(1). Let the ordinary execution cost for a supported instruction i be:
T_std,i = D_std,i + I_i + S_i
where D_std,i is ordinary decode cost, I_i is indirect address-formation cost, and S_i is the instruction semantics after the target address is known. The quickened cost is:
T_q,i = D_q,i + S_i
The conditional per-execution gain is therefore:
G_i = T_std,i - T_q,i = I_i + D_std,i - D_q,i
Equation (6) makes no assumption about a particular flash technology or candidate applet. The transformation is beneficial whenever the removed indirect work exceeds any additional direct-decoder cost.

3.6. Memory-Access and Branch Decomposition

For an implementation in which I_i consists of d_i memory dereferences with costs L_i,1 through L_i,d_i, plus c_i processor cycles of address-selection work at frequency f, the gain can be expressed as:
G_i = SUM(k=1..d_i) L_i,k + c_i/f + D_std,i - D_q,i
The quantities d_i, L_i,k, c_i, and f are platform parameters. The model therefore does not claim a universal number of eliminated reads, a universal flash latency, or a universal percentage speedup.

3.7. Transaction-Level Bound

Let n_i be the number of executions of supported instruction class i during a transaction. The total analytical gain is:
Delta T = SUM_i n_i x G_i
If a platform establishes a conservative lower bound G_min > 0 for all quickened instructions on the relevant path and N = SUM_i n_i, then:
Delta T >= N x G_min
Equation (9) is the appropriate point at which a future candidate applet’s instruction count may be substituted. Without a disclosed applet or trace, N remains a variable rather than a claimed typical value.

3.8. Installation and Space Cost

Let m be the number of reference locations considered for quickening. The added linker work is O(m). Because the ordinary linker already visits those locations, the practical increment is a constant amount of address encoding and byte replacement per accepted location. The Method component incurs zero byte growth. The interpreter incurs additional decoder logic for the private families; its exact code-size cost is implementation-specific. The encoding itself requires no extra per-instruction runtime metadata.

3.9. Worked Analytical Example

Assume an invokestatic reference is accepted by the ordinary linker and resolves to logical address a = 0x31A72. The address lies below 0x80000 and therefore fits in 19 bits. For the proposed invokestatic base B = 0xE0:
Preprints 231084 i003
At execution, the interpreter computes family = 0xE3 AND 0xF8 = 0xE0, recognizes invokestatic, and reconstructs:
((0xE3 AND 0x07) << 16) OR (0x1A << 8) OR 0x72 = 0x31A72
The example demonstrates encoding and decoding only. It is not a timing experiment and does not imply a particular application workload.

4. Discussion

4.1. CAP-File Input and Interpreter Compatibility

The proposal accepts an ordinary CAP file and does not require application source changes or a new off-card compiler format. Quickened bytes are produced only inside the platform-controlled installed representation. The appropriate compatibility claim is therefore CAP-file input compatibility and Method-layout preservation, not universal binary compatibility of the private installed image across unmodified JCVMs.
An unmodified interpreter cannot execute the private opcodes. The target JCVM must add decoding and execution support and retain the ordinary installed form as a fallback for out-of-range, movable, unsupported, or dynamically resolved targets. Mixed ordinary and quickened execution is therefore possible within one platform.

4.2. Access Control and Applet Isolation

Quickening occurs after ordinary installation checks and must not bypass package visibility, class access, method access, field access, type compatibility, firewall, context, or transaction rules. If the standard platform performs an execution-time security check that is not solely part of address formation, that check remains in the quickened handler. The optimization removes indirection, not security semantics.

4.3. Address Stability and Lifecycle Operations

Absolute addresses are valid only while the target layout is stable. Offline ROM-mask or flash-image generation naturally satisfies this condition. Post-issuance installation can also satisfy it when installed code and static storage are non-moving. A platform that compacts, relocates, replaces, or unloads targets must either re-link every dependent quickened instruction, maintain a stable addressing abstraction, or avoid quickening those references.

4.4. Applicability

The scheme is relevant to JCVM deployments in which supported targets remain stable after linking and the ordinary installed representation still performs non-trivial address-formation work during execution. Potential domains include payment, transit, access control, UICC/SIM services, and embedded authentication. UICC Java Card environments are standardized through specifications such as ETSI TS 131 130 [6], but the analytical result is independent of any one domain.
The design is particularly natural during offline ROM-mask or flash-image generation, where the full memory map is known before personalization. It may also be used during on-card installation when the platform guarantees target stability. The 512 KiB range is a property of the encoding and is not a claim that every platform allocates exactly that amount to Java Card code or data.

4.5. Limitations and Threats to Validity

The principal limitations are:
  • Implementation scope: a JCVM that already uses direct addresses or an equally efficient quickened representation may obtain no benefit.
  • Opcode availability: the proposed hexadecimal mapping must be checked against the precise specification edition and all private opcodes in the target implementation.
  • Address range: targets outside the 19-bit logical domain require fallback, segmentation, another encoding, or an additional translation level.
  • Relocation: movement of target methods or static storage invalidates direct addresses unless dependent instructions are re-linked.
  • Dynamic dispatch: invokevirtual and invokeinterface are not covered.
  • Decoder overhead: a direct handler may have a different decode cost; the analytical model includes this term rather than assuming it is zero.
  • Security behavior: execution-time checks unrelated to reference address formation cannot be removed.
  • Evidence boundary: the study proves structural and conditional analytical properties but reports no measured end-to-end speedup.

4.6. Interpretation

The proposal is best understood as Java Card-specific bytecode quickening. It shifts a repeated representation-conversion step from execution to installation while preserving instruction length. Its novelty is not the general concept of early resolution, but the combination of an in-place three-byte encoding, a 19-bit direct logical address, selected Java Card instruction classes, and explicit preservation of Method-component byte positions.
The theoretical treatment clarifies what can and cannot be concluded without an applet. Correctness, encoding capacity, layout invariance, transformation complexity, and conditional cost reduction can be established without a benchmark. A numerical application-level speedup cannot. For each optimized execution, the scheme removes the baseline implementation’s indirect address-formation work, subject to decoder cost and the stated assumptions.
This formulation is falsifiable. An implementer can examine a target JCVM, identify I_i, count or measure its components, and substitute those values into Equations (6)-(9). A system with I_i approximately zero predicts negligible benefit; a system with non-volatile-memory table dereferences predicts a larger benefit. Both outcomes are consistent with the model.

5. Conclusions

This paper presented a link-time quickening scheme for selected Java Card method-invocation and static-field instructions. The scheme encodes a 19-bit direct logical address in an implementation-private opcode family and the instruction’s two existing operand bytes. It preserves the three-byte instruction length and therefore leaves subsequent Method-component offsets unchanged.
Under explicit linker-correctness, address-stability, decoder-correctness, address-range, and opcode-separation assumptions, the quickened form reconstructs the same target and preserves the observable instruction semantics. The analytical model shows that the per-execution gain is the removed indirect address-formation cost adjusted for any decoder-cost difference. No candidate applet or hardware benchmark is required to establish these structural results, and no universal speedup is claimed.
The scheme provides a practical design option for Java Card platforms that retain indirect linked references after installation. It can be introduced incrementally with fallback to the ordinary representation and without changing application source code, CAP-file input, or Method-component length. Future work may include a machine-checked proof, a reference implementation, verified opcode allocation for specific Java Card editions, and optional empirical evaluation on disclosed workloads.

Author Contributions

Both authors contributed to conceptualization, methodology, formal analysis, and writing. Alp Sardağ supervised the study and coordinated the technical review. All authors have read and agreed to the published version of the manuscript.

Funding

This research received no external funding.

Data Availability Statement

No new data were created or analyzed in this theoretical study. Data sharing is not applicable to this article.

Acknowledgments

During the preparation of this manuscript, the authors used OpenAI ChatGPT to assist with language editing, manuscript structuring, and formatting. The authors reviewed and edited the output and take full responsibility for the content of the publication.

Conflicts of Interest

The authors declare no conflicts of interest.

Abbreviations

CAP, Converted Applet; JCVM, Java Card Virtual Machine; UICC, Universal Integrated Circuit Card.

Appendix A

Appendix A.1. Transformation Summary

Table A1. Transformation stages, quickening actions, and fallback behavior.
Table A1. Transformation stages, quickening actions, and fallback behavior.
Stage Ordinary action Additional quickening action Failure behavior
CAP input Parse standard CAP components. None. Reject malformed input as usual.
Validation Resolve and validate reference. None. Reject invalid or inaccessible reference as usual.
Layout Determine final target address a. Check a < 2^19 and address stability. Retain ordinary linked form.
Rewrite Produce installed representation. Overwrite three bytes with q, b1, b2. Retain ordinary linked form.
Execution Form target through baseline representation. Decode private family and reconstruct a. Ordinary handler remains available.

Appendix B

Appendix B.1. Claim Boundaries

Table A2. Supported conclusions and conclusions not made in this study.
Table A2. Supported conclusions and conclusions not made in this study.
Supported conclusion Conclusion not made
The encoding covers 512 KiB and is exactly reversible. Every Java Card platform exposes a 512 KiB Java addressing domain.
The rewrite preserves instruction and Method-component length. An unmodified JCVM can execute the private opcodes.
Semantics are preserved under the stated assumptions. All runtime security checks can be removed.
The scheme removes the baseline indirect address-formation work. Every compliant JCVM performs the same number of flash reads.
Transaction gain is SUM_i n_i x G_i. A typical EMV or SIM transaction saves a fixed number of milliseconds.
A candidate applet is unnecessary for the formal construction. Application-level speedup can be claimed without a disclosed workload.

References

  1. Oracle. Java Card 3 Platform Virtual Machine Specification, Classic Edition, Version 3.0.5; Oracle: Austin, TX, USA, 2015; Available online: https://docs.oracle.com/javacard/3.0.5/JCVMS/JCVMS.pdf (accessed on 26 July 2026).
  2. Bouffard, G.; Giraud, V.; Gaspard, L. Java Card Virtual Machine Memory Organization: A Design Proposal. arXiv 2021, arXiv:2110.10037. [Google Scholar] [CrossRef]
  3. Czajkowski, G.; Daynès, L.; Nystrom, N. Code Sharing among Virtual Machines. In Proceedings of the ECOOP 2002 - Object-Oriented Programming Lecture Notes in Computer Science, Málaga, Spain, 10-14 June 2002; Springer: Berlin/Heidelberg, Germany, 2002; Volume 2374, pp. 155–177. [Google Scholar] [CrossRef] [PubMed]
  4. Lounas, R.; Mezghiche, M.; Lanet, J.-L. Towards a General Framework for Formal Reasoning about Java Bytecode Transformation. Electron. Proc. Theor. Comput. Sci. 2013, 122, 63–73. [Google Scholar] [CrossRef]
  5. Mesbah, A.; Lanet, J.-L.; Mezghiche, M. Reverse Engineering Java Card and Vulnerability Exploitation: A Shortcut to ROM. Int. J. Inf. Secur. 2019, 18, 85–100. [Google Scholar] [CrossRef]
  6. ETSI. ETSI TS 131 130 V15.3.0; Digital Cellular Telecommunications System (Phase 2+); Universal Mobile Telecommunications System; LTE; 5G; (U)SIM Application Programming Interface; (U)SIM API for Java Card. European Telecommunications Standards Institute: Sophia Antipolis, France, 2020. Available online: https://www.etsi.org/deliver/etsi_ts/131100_131199/131130/15.03.00_60/ts_131130v150300p.pdf (accessed on 26 July 2026).
Table 1. Supported instruction groups and link-time resolvability.
Table 1. Supported instruction groups and link-time resolvability.
Semantic group Source instruction(s) Reason the target is link-time resolvable
Static method call invokestatic The linked symbolic reference selects one static method.
Special method call invokespecial The special target is fixed by the accepted class relation and linked instruction semantics.
Static read getstatic_b, getstatic_s, getstatic_a, getstatic_i The static storage location is fixed after linking and layout.
Static write putstatic_b, putstatic_s, putstatic_a, putstatic_i The static storage location is fixed after linking and layout.
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.