Computer Science and Mathematics

Sort by

Article
Computer Science and Mathematics
Data Structures, Algorithms and Complexity

Gaurav Sablok

Abstract: Chloroplast genomes of land plants typically possess a conserved quadripartite structure consisting of large and small single-copy regions (LSC and SSC) separated by two large inverted repeats (IRs). Here, we present chloromapper, a lightweight Rust program designed to automate the arrangement of pre-assembled chloroplast contigs into a complete circular genome. Given three or more plastome contigs in arbitrary order, orientation, and fragmentation, chloromapper identifies IR sequences using inverted-duplicate or read-depth evidence, chains fragmented single-copy contigs based on exact end overlaps, assigns sequences to the LSC and SSC regions, and resolves their relative orientations using IR–single-copy junctions. The program subsequently reconstructs the canonical circular LSC–IRb–SSC–IRa arrangement, merging overlapping contigs and representing unresolved junctions or pre-existing assembly gaps with ambiguous nucleotides.We evaluated chloromapper using synthetic assemblies derived from the complete chloroplast genome of Arabidopsis halleri (GenBank KX886356.1), including resolved-IR, collapsed-IR, and fragmented short-read-style scenarios.

Article
Computer Science and Mathematics
Data Structures, Algorithms and Complexity

Giulio Ruffini

,

Francesca Castaldo

Abstract: Scientific discovery seeks regularities that support explanation and prediction. Compression makes their reuse explicit: shared structure is described once, while parameters specify individual cases. A model can therefore pay for itself through repeated use even when its description is not minimal, provided it captures reusable structure in the data. Because regularities are often easier to identify in simpler systems, a reductionist approach is often employed: discover the laws of the parts and treat them as fundamental. Yet knowing those laws does not by itself provide useful coarse-grained models of the larger systems they compose. Here we formulate Anderson’s distinction between reduction and construction for finite algorithmic observers and prove three barriers to such construction. First, an observer’s coarse-grained record may retain so much information about initial or boundary conditions that no substantially shorter description exists, even when the underlying laws are simple and known. Second, when a shorter description does exist, open-ended search can eventually find one, but there is no computable bound on how long this may take, and no algorithm that always halts with a valid description can succeed in every case. Third, every such algorithm has blind spots at all sufficiently large lengths: records it leaves unshortened even though they admit descriptions of only logarithmic length. Yet observers do discover useful models. We call the event in which an observer acquires a representation relevant to its objective and a reusable model that reveals previously unavailable regularity algorithmic emergence. We give a sufficient certificate in explicit code lengths: the pair passes when, with its own description cost included, it compresses the retained data relative to an agreed baseline and yields further savings on later observations. Even for a stream that repeats one block, no computable procedure that always halts with valid codes can guarantee finding a passing pair with the observations available whenever one exists, nor remain within a fixed number of bits of the best qualifying complete code. Favorable structure, including symmetry and restricted model classes, can nevertheless make discovery feasible.

Article
Computer Science and Mathematics
Data Structures, Algorithms and Complexity

Gaurav Sablok

Abstract: Circular ideograms are used to visualize bacterial genome, synteny and genomic conversation. I describe an angle-allocation algorithm that unifies single-genome ideogram layout and multi-genome synteny layout as one recursive construction, a CIGAR-exact depth-binning algorithm operating on a pure-Rust BAM decoder, and a minimal SVG primitive set—the annular sector and the through-center ribbon—sufficient to reproduce Circos’s core visual vocabulary. I implement this as bacircos, a dependency-light CLI, and validate the synteny algorithm against a real minimap2 alignment of two synthetic strains related by a seeded inversion and deletion: the tool recovers both as direct, non-heuristic consequences of the layout and ribbon construction.

Article
Computer Science and Mathematics
Data Structures, Algorithms and Complexity

Feng Li

,

Yali Si

,

Zijun Yan

,

Xin Li

,

Jiacheng Wang

,

Yongchao Jiang

Abstract: Point-of-interest (POI) recommendation can provide personalized and intelligent location recommendation services, which is of significant value in location-based social networks. However, current research lacks comprehensive awareness of user mobile contexts, and POI recommendation algorithms perform calculations over all check-in data, resulting in high computational complexity and low accuracy. To address these issues, we propose a lightweight POI recommendation method called DF-LR, which combines mobile direction awareness and filtering mechanisms. Specifically, for mobile context awareness, we obtain the user’s location via GPS and the movement direction via the compass sensor in smartphones, then reduce noise using the median filtering method and propose a direction similarity calculation method. Additionally, we devise two filtering mechanisms to achieve lightweighting: POI filtering is performed by mining the distance features of user’s adjacent visit positions, and check-in data filtering is performed by mining the features of users’ check-ins at the same position. Finally, Jaccard-based user similarity and movement direction correlation are combined to calculate recommendation probability. Extensive performance evaluation experiments show that DF-LR improves precision, recall and F1-measure with lower computation overhead compared to baseline POI recommendations.

Technical Note
Computer Science and Mathematics
Data Structures, Algorithms and Complexity

Usama Mehboob

Abstract: Anonymizing sensitive healthcare data while preserving data utility is always a tradeoff between suppression and generalization. In this experiment, we employ a genetic algorithm to search the anonymization policy space using synthetic healthcare-style data. Each policy candidate specifies different levels of generalization for quasi-identifiers such as age, a five-digit numeric location code, and sex, while any equivalence class that does not satisfy the constraints of k-anonymity or distinct ℓ-diversity is suppressed. The GA experiment is carried out under constraints of k = 5, ℓ = 2, and a 30% suppression limit. There are 32 possible anonymization policies, and the search space is intentionally kept small so that the results of the GA-based search can be compared directly with exhaustive search to identify its accuracy, gaps, and limitations. Across 60 seeded runs using three different table sizes, the GA search recovered the exact Pareto set in 59 runs while evaluating a median of 28–29 policies. These results show that the GA was capable of reliably recovering the Pareto-optimal solutions in this small synthetic setting, but GA still ended up evaluating most of the policies. The efficiency advantage of GA has yet to be evaluated in the future on large real world healthcare data with substantially larger policy spaces to establish if GA indeed offers more clinical utility with less compute compared to other traditional methods.

Article
Computer Science and Mathematics
Data Structures, Algorithms and Complexity

Saeid Jafari

,

Paulraj Gnanachandra

,

Giorgio Nordo

,

Lorenzo Affè

,

Florentin Smarandache

Abstract: We introduce a neutrosophic framework for studying randomized rounding methods for the Traveling Salesman Problem (TSP). The approach is motivated by the work of Gharan, Saberi, and Singh on randomized rounding for the graphic TSP. In the proposed model, each edge is endowed with a cost interval and a neutrosophic assessment consisting of truth, indeterminacy, and falsity degrees. A scalarization rule converts this information into an effective edge cost, allowing classical linear programming relaxations and combinatorial optimization methods to be applied. We formulate a neutrosophic subtour relaxation, discuss randomized spanning-tree selection, and describe the role of minimum-cost T-joins in constructing tours. We also establish elementary properties of the scalarized model and give a conditional approximation result. The conditions needed for such a guarantee are stated explicitly; in particular, the classical approximation bound does not follow from neutrosophic notation alone. The paper provides a mathematical starting point for incorporating uncertain, incomplete, and conflicting edge-cost information into randomized approximation algorithms for the TSP.

Article
Computer Science and Mathematics
Data Structures, Algorithms and Complexity

Frank Vega

Abstract: We explore the complexity-theoretic scenario in which \(\mathrm{P} = \mathrm{NP}\) but \(\#\mathrm{P} \neq \mathrm{FP}\), and ask what it would imply for the Birch–Swinnerton-Dyer conjecture (BSD) for the congruent number curves \(E_n: y^2 = x^3 - n^2x\). Neither hypothesis of the scenario is known to hold; the aim is to determine rigorously what follows if both do. The link is Tunnell's theorem, which expresses the congruence of \(n\) through the representation counts \(C_n\) and \(D_n\) of \(n\) by the ternary forms \(8x^2 + 2y^2 + 64z^2\) and \(8x^2 + 2y^2 + 16z^2\). We prove four unconditional results. First, for square-free \(n \equiv 2 \pmod 8\) one has \(D_n = 2h(-4n)\), where \(h(-4n)\) is the class number of \(\mathbb{Q}(\sqrt{-n})\), while \(D_n = C_n = 0\) for square-free \(n \equiv 6 \pmod 8\). Second, class numbers of imaginary quadratic orders are computable in \(\mathrm{FP}^{\Sigma_2^{\mathrm{P}}}\), hence in \(\mathrm{FP}\) if \(\mathrm{P} = \mathrm{NP}\). Third, the difference \(C_n - h(-4n)\) is the \(n/2\)-th Fourier coefficient of Tunnell's weight-\(3/2\) cusp form, so that, by Tunnell's \(L\)-value formula and Ono's evaluation of the BSD prediction, BSD for \(E_n\) gives \(C_n = h(-4n) \pm \tau(n/2)\sqrt{|\mathrm{Sha}(E_n)|}\) in rank zero and \(C_n = h(-4n)\) in positive rank. Fourth, any exact counting reduction from \(\#\mathrm{SAT}\) to \(D_n\) would collapse the polynomial hierarchy to \(\Delta_3^{\mathrm{P}}\). Our main theorem combines these facts with the enumerative-counting theorem of Cai and Hemachandra: if \(\mathrm{P} = \mathrm{NP}\) and \(\#\mathrm{P} \neq \mathrm{FP}\), and if (i) \(\#\mathrm{SAT}\) reduces exactly to \(C_n\) on square-free \(n \equiv 2 \pmod 8\) (the Cusp Reduction Conjecture) and (ii) positivity of the rank of \(E_n\) and, in rank zero, the order of the Tate--Shafarevich group of \(E_n\) are polynomial-time computable with an NP oracle (the Sha-Tractability Hypothesis), then BSD fails for infinitely many of the curves \(E_n\); since Rubin's theorems give the finiteness of \(\mathrm{Sha}(E_n)\) and the odd part of the BSD formula whenever \(L(E_n,1) \neq 0\), the failure must lie in the rank statement or in the \(2\)-part of the rank-zero formula. We make explicit that the conclusion is conditional on (i) and (ii), and that, logically, the scenario forces the failure of at least one of BSD, (i) and (ii).

Article
Computer Science and Mathematics
Data Structures, Algorithms and Complexity

Frank Vega

Abstract: The Minimum Vertex Cover problem is NP-hard. Its classical polynomial-time approximation ratio is 2, and Khot and Regev showed that, assuming the Unique Games Conjecture (UGC), no polynomial-time algorithm can achieve a factor \(2-\varepsilon\) for any fixed \(\varepsilon>0\). We present FindVertexCover, a linear-time ensemble algorithm that runs seven independently valid vertex-cover heuristics on the input graph \(G=(V,E)\) and returns the smallest cover found, then prunes redundant vertices. Four of the seven candidates—and the final redundant-vertex pruning pass applied to each of them—adapt a published ensemble that proves an unconditional \(O(n+m)\) time and space bound and a worst-case approximation ratio at most 2 for every graph, where \(n=|V|\) and \(m=|E|\). \textsc{Salvador} extends that ensemble with two further linear-time candidates of its own and a sixth candidate, SolveVC, that reduces \(G\) to a linear-size planar forest core via a weighted Minimum Independent Dominating Set (MIDS) gadget and solves that gadget with an accuracy-controlled Baker-style PTAS whose layering width is \(k=\lceil 1/\varepsilon\rceil\); non-core edges are covered by a greedy repair step. At the package default \(\varepsilon=1\), the Baker layering width is \(k=1\) and the PTAS pass degenerates to its linear-time greedy baseline, so every one of the seven candidates—and hence the ensemble as a whole—runs in worst-case \(O(n+m)\) time; because one candidate is the classical maximal-matching heuristic, the ensemble unconditionally achieves approximation ratio at most 2 on every graph, a theorem rather than a conjecture. We conjecture a universal ratio bound \(R^\star<2\) for the default call; on a reproducible suite of 1,746 test graphs, ranging from exhaustively enumerated small graphs to random and adversarial graphs with up to 200,000 vertices, the observed ratio never exceeds $1.61$, well below the classical \(7/4\) mark. If a bound below 2 were proved for the ensemble, then, under the standard assumption \(\mathrm{P}\neq\mathrm{NP}\), the Khot–Regev UGC-based hardness theorem for Vertex Cover would force the Unique Games Conjecture to be false. An open-source implementation is released as the Salvador package (v0.0.7).

Article
Computer Science and Mathematics
Data Structures, Algorithms and Complexity

Giulio Ruffini

Abstract: The idea that successful regulation requires an internal world model helps motivate both the generative models of active inference and the modeling engine of Kolmogorov Theory's algorithmic agent. Conant and Ashby's good regulator theorem and the internal model principle support this idea under specific assumptions. We ask when successful regulation implies that a regulator already contains information about the world. We compare finite output records of the same world with and without regulation, holding initial world conditions, disturbance inputs, and measurement rules fixed, and quantify the complexity reduction using prefix Kolmogorov complexity. Our starting point is precise information accounting, organized through an information-conserving logical construction. A computable reversible realization preserves the descriptions needed to recover the initial world and reconstruct its uncontrolled output. This realization may be auxiliary: physical determinism and reversibility are not general requirements of the resulting finite-record bounds. The Good Algorithmic Regulator Theorem (GART) bounds the reduction by the mutual algorithmic information initially shared by world and regulator plus a counterfactual residual, up to logarithmic coding overhead. The residual is the additional description needed to reconstruct the uncontrolled output given the regulated output and the regulator's initial description. A large reduction therefore requires substantial initial shared information when the residual is independently bounded well below the reduction. When actions, final regulator memory, and complementary world records suffice for reconstruction, they account for the residual: model it, transmit it, or leave it in the world. A generative-model example makes the output constant while preserving reconstruction information; feedback and clamps can also yield large reductions, with information carried by actions or retained in the world. Compact intervention knowledge can sustain regulation while supplying little additional compression of the uncontrolled record. The deterministic bound applies to individual finite records without a disturbance distribution or a prior over programs. The earlier Algorithmic Regulator Theorem (ART) bounded the posterior weight of a specified world–regulator explanation, but a fixed-clamp family with nonvanishing posterior mass defeats its unrestricted aggregate inference. Conditioning on a bounded residual guarantees, under any probability law on admissible world–regulator pairs, that initial shared information cannot fall below the reduction by more than the residual bound and coding allowance. For independent initial draws from fixed computable laws, an exponential bound limits the probability of a large reduction with a small residual.

Article
Computer Science and Mathematics
Data Structures, Algorithms and Complexity

Frank Vega

Abstract: The maximum independent set problem asks for the largest set of pairwise non-adjacent vertices in an undirected graph and is NP-hard in general. This paper describes Esperanza, an approximation algorithm built around the linear-time Hvala vertex cover approximation. It establishes a first candidate independent set from the complement of Hvala's cover and evaluates a dynamic minimum-degree Caro-Wei baseline that itself runs in linear time. For every vertex in the cover---and explicitly including the graph's maximum degree vertex---a repair step removes it, reinstates its neighbors, and regrows a maximal independent set using an \( O(n + m) \) conflict resolution and bucket-sort degree-scan phase, keeping the best of all these attempts. A further phase searches for large independent sets by extracting maximal cliques of size at least two from the complement graph via the clique-preserving disjoint-set structure FastCliqueUF; since a single vertex's complement neighbourhood can fragment into as many as \( \Theta(n) \) disjoint clique components in the worst case, each requiring its own linear-time repair-and-grow call, the algorithm sorts the candidates by size and processes only the \( \max(2,\lfloor\log_2 n\rfloor) \) largest per triggering vertex, on the reasoning that the largest available clique is the one most likely to improve the incumbent. This bounds the number of such calls per vertex by \( O(\log n) \), giving an overall worst-case running time of \( O(n^3\log n) \). We prove that the output is always a maximal independent set and that the Caro-Wei baseline together with the explicit \( v_{\max} \) repair yields a structural approximation ratio of \( O(\sqrt{n}) \), unaffected by the cap. The complement-clique phase can only improve the concrete solution and never degrades the proven ratio. Empirical evaluation on 30\,000 MILP-certified instances yields a worst-case ratio of 1.20 and a global mean ratio of 1.000018; a further diagnostic on a larger adversarial instance shows the cap binding on half of all triggering vertices there, doing substantial, measurable work.

Article
Computer Science and Mathematics
Data Structures, Algorithms and Complexity

Frank Vega

Abstract:

We present AEGYPTI (v0.5.6), a combinatorial triangle-detection framework for an undirected simple graph \(G=(V,E)\) with \(n=|V|\) vertices and \(m=|E|\) edges. Execution is dynamically dispatched by density at the structural threshold \(\lceil n^{4/3}\rceil\). When \(m\le\lceil n^{4/3}\rceil\), an optimized Chiba–Nishizeki adjacency-intersection routine, sorted by non-decreasing degree, is invoked. For denser instances, the algorithm employs a randomized square-root partitioning strategy that isolates induced subgraphs and geometrically prunes intra-partition edges in \(\mathcal{O}(n^{2})\) time per iteration. Coupled with a linear-time Caro–Wei independent-set computation on the complement (and bipartite short-circuits), this strategy yields an expected runtime of \(\mathcal{O}(n^{2.5}\log n)\) for dense regimes. The randomness is drawn from a cryptographically secure source, guaranteeing a true Las Vegas algorithm whose expectation is taken solely over internal coin flips. We prove soundness, completeness, and the stated complexity bound in full detail. The resulting subcubic bound constitutes a randomized (Las Vegas) combinatorial counter-example to the generalized assumptions of the Combinatorial Boolean Matrix Multiplication (BMM) Conjecture. A public reference implementation is available in the aegypti Python package.

Article
Computer Science and Mathematics
Data Structures, Algorithms and Complexity

Parth Kumar Sinha

Abstract: The in-memory write buffer of an LSM-tree key-value store is searched on every read and written on every insert, so its cost per operation sets a floor on the whole engine. RocksDB uses a concurrent skip list, which chases pointers across independently allocated nodes and evaluates one data-dependent branch per key comparison. We present Aparajita, a MemTable representation that replaces the skip list with a list of cache-line-sized nodes, each holding fifteen 32-bit order-preserving key surrogates and a sentinel in one line, searched by a branchless SIMD kernel. Three design decisions carry the result. A relational vector compare over a sorted node yields a mask whose population count is the lower bound directly, so ordered search costs one compare, one movemask and one popcount with no branch. Surrogates are taken after the node’s shared prefix rather than from the start of the key, without which an absolute surrogate takes one value across all 200,000 keys in five of eight realistic distributions, including the keyspace this paper’s own evaluation runs on. Nodes are append-only, and the sorted order over their slots is a 64-bit word, so an insert is two stores into a free slot followed by one release store that publishes them. We implement Aparajita as a RocksDB plugin selectable by name without patching RocksDB sources, and evaluate it against the default skip list and VectorRep on a 12-core Emerald Rapids host. Point lookups over a resident MemTable are 22% to 39% faster than the skip list at 1, 4, 16 and 64 threads, non-overlapping across five runs per configuration at every point but one, backed by 55% fewer retired instructions and 40% fewer L1 misses per lookup. Ordered seeks are 15% to 29% faster, but a seek followed by ten iterator steps is 3% to 4% slower, and the representation charges 1.4 times the skip list’s arena per key. The skip list is 48% to 52% slower on insert at the representation level, and Aparajita is 18.8% faster in single-threaded db_bench, but the multi-threaded db_bench write path is bounded by RocksDB’s write group rather than by the MemTable: an insert there retires over 22,000 instructions in both representations. We report that ceiling rather than a write scaling claim the data does not support.

Article
Computer Science and Mathematics
Data Structures, Algorithms and Complexity

Zhao Song

Abstract: We prove two deterministic inapproximability results. First, for every fixed \(0<\epsilon<1/2\), Euclidean \(\mathrm{GapCVP}^{(2)}\) is NP-hard with gap factor \(n^{1/2-\epsilon}\) under deterministic polynomial-time many-one reductions, where \(n\) denotes the lattice rank. Consequently, the Euclidean closest vector problem is NP-hard to approximate within the same factor. This improves the previous \(n^{1/400}\) hardness factor in Chapter 7 of the OpenAI report [Ope25]. Aharonov and Regev gave short certificates for both the YES and the NO case of \(\operatorname{GapCVP}^{(2)}\) at gap factor \( C\sqrt n \), placing that problem in \(\mathrm{NP}\cap\mathrm{coNP}\) for an absolute constant \(C>0\) [AR05]. An NP-hard problem lying in \(\mathrm{coNP}\) would give \(\mathrm{NP}=\mathrm{coNP}\), so the factor \(n^{1/2-\epsilon}\) above cannot be improved to \( C\sqrt n \) unless the two classes coincide. Second, for every fixed \(0<\epsilon<1\), the gap versions of binary nearest codeword and binary syndrome decoding are NP-hard with factor \(n^{1-\epsilon}\) under deterministic polynomial-time many-one reductions, where \(n\) denotes the binary block length. Consequently, both optimization problems are NP-hard to approximate within the same factor. This improves the previous \(n^{1/200}\) hardness factor in Chapter 7 of the OpenAI report [Ope26].

Article
Computer Science and Mathematics
Data Structures, Algorithms and Complexity

Alp Sardağ

,

Abdullah Sardağ

Abstract: Java Card virtual machines operate in resource-constrained secure elements where repeated reference indirection can add latency to frequently executed instructions. This paper proposes a link-time bytecode quickening scheme for statically resolvable method-invocation and static-field instructions. After ordinary installation-time validation and symbolic resolution, the linker rewrites each selected three-byte instruction into an implementation-private direct-address form of the same length. The low three bits of an aligned opcode family encode the three most significant address bits, while the existing two operand bytes encode the remaining sixteen bits, yielding a 19-bit logical address space of 512 KiB. The study is formal and analytical rather than experimental. It defines the transformation and decoder, proves address round-trip correctness, establishes preservation of Method-component length and control-flow offsets, and gives a conditional semantic-equivalence argument under explicit linker-correctness, address-stability, decoder-correctness, address-range, and opcode-separation assumptions. A parametric execution-cost model shows that the gain equals the removed indirect address-formation work, adjusted for any decoder-cost difference. The proposal targets JCVM implementations that retain indirect post-link reference representations; dynamic dispatch instructions are excluded. The scheme preserves ordinary CAP-file input and application source code while requiring a modified linker and interpreter for the installed representation.

Article
Computer Science and Mathematics
Data Structures, Algorithms and Complexity

Zijian Zeng

Abstract: We settle the complexity of reaching a globally Pareto-efficient allocation by mutually improving swaps along the edges of a tree. The problem is NP-complete, even though reachability of a specified full allocation is polynomial-time decidable on trees. The proof is a linear-size reduction from the top-target restriction of object reachability on trees. Its two devices are a fixed sentinel, which certifies global Pareto inefficiency whenever the target object has not arrived, and private cleanup leaves, which turn every successful target-reaching state into a globally Pareto-efficient allocation. The construction doubles the number of agents, preserves the tree property, and increases maximum degree by at most one. We also separate the standard global notion used here from Pareto maximality within the reachable set, which is the notion used in the earlier social-network allocation literature.

Article
Computer Science and Mathematics
Data Structures, Algorithms and Complexity

Luigi Laura

,

Marco Parrillo

,

Alessio Pascucci

,

Marco Rossi

,

Valerio Rughetti

Abstract: Quantum algorithms have been proposed for most computational tasks in trading, but the resulting literature is not a single body of work with a single standard of evidence. This survey partitions it into strands with distinct foundations and assesses each against the strongest available evidence. Pricing and market-risk computation rest on quantum amplitude estimation and carry a proven quadratic advantage in inverse precision; circuit-level resource estimates place the requirement at roughly 8000 logical qubits and a logical clock rate of tens of megahertz, several orders of magnitude beyond demonstrated hardware, with joint estimation of prices and sensitivities the main lever that lowers it. Combinatorial formulations of portfolio selection, multi-period trading trajectories and arbitrage detection map onto quadratic unconstrained binary optimization and have the field’s most mature hardware demonstrations, but an extensive benchmark on 250 real-stock instances finds classical mixed-integer programming and tailored heuristics dominant. Quantum machine learning for return prediction has been substantially dequantized, and a controlled evaluation on cross-sectional equity returns attributes reported advantages to evaluation protocol rather than to the model. A fourth strand exploits entanglement for decision coordination under latency constraints, where the advantage follows from Bell’s theorem rather than a complexity assumption. We identify the classical-quantum interface, superquadratic structure, resource-estimate methodology, nonlocal-game design and adversarial benchmarking as the productive open directions.

Article
Computer Science and Mathematics
Data Structures, Algorithms and Complexity

Rodolfo Bojorque

,

David Yánez-Peter

,

Miguel Arcos

Abstract: Recommender systems increasingly incorporate graph embeddings and graph neural networks to capture high-order relationships between users and items. However, the additional complexity of these approaches does not necessarily guarantee better recommendation quality than strong classical and latent-factor baselines. This study presents a reproducible comparison of six recommendation models representing four methodological families: Logistic Regression, Random Forest, Matrix Factorization with Bayesian Personalized Ranking, DeepWalk, node2vec, and LightGCN. The experiments were conducted on the MovieLens 1M dataset using a per-user temporal split. For each user, the most recent positive interaction was assigned to testing, the preceding interaction to validation, and all earlier positive interactions to training. All models were evaluated using identical candidate sets containing one held-out positive movie and 99 sampled unobserved movies. Performance was measured using Recall, Precision, Hit Rate, and NDCG at multiple cutoffs, complemented by bootstrap confidence intervals, paired statistical tests, computational-efficiency measurements, and analyses by user activity and movie popularity. Matrix Factorization achieved the best overall performance, reaching a Recall@10 of 0.7458 and an NDCG@10 of 0.4558, representing an approximately 56% improvement in NDCG@10 over Random Forest, the strongest classical baseline. LightGCN did not significantly outperform Logistic Regression and remained below Random Forest, despite its higher computational cost. DeepWalk and node2vec obtained similar and substantially lower aggregate results. Popularity-based analysis further revealed that classical models and LightGCN achieved substantially higher ranking effectiveness for popular movies, whereas Matrix Factorization maintained comparatively stronger performance for less-popular items. These findings demonstrate that model complexity alone is not a reliable indicator of recommendation effectiveness and highlight the importance of strong baselines, standardized evaluation, and reproducible experimental protocols.

Article
Computer Science and Mathematics
Data Structures, Algorithms and Complexity

Zhao Song

Abstract: We prove that every finite two-player game \(G\) with entangled value \(\omega^*(G)=1-\epsilon\) satisfies \[ \omega^*(G^{\otimes n}) \le\exp(-\Omega(\frac{\epsilon^3}{\epsilon+ \ell }n)) \] for every $n\ge1$, where \(\ell:=\log(|A| |B|)\), and \(A\) and \(B\) are the answer alphabets. Compared with Chapter~6 of the OpenAI report [Ope26], this improves the gap exponent from thirteen to three and matches the cubic gap dependence in Holenstein's general classical bound [Hol09]: \[ \omega(G^{\otimes n})\le\exp(-\Omega( \frac{ (1-\omega(G))^3}{1+\ell} n)). \] The proof replaces the randomly shifted logarithmic grid used in quantum correlated sampling by smooth soft labels. This makes the relevant label infidelity quadratic in the distance between state descriptions and avoids a Jensen loss when averaging over questions. Together with the postselection argument, these improvements yield the cubic gap dependence stated above.

Article
Computer Science and Mathematics
Data Structures, Algorithms and Complexity

Szymon Łukaszyk

,

Piotr Masierak

Abstract: Assembly theory measures a string by the shortest history that builds it. We sharpen this picture in two directions. First, we attach a discrete Dirichlet energy to an assembly space and show that it splits into the size of the space and a secondary energy charging each step vertex the square of its secondary (pathway depth) increment. The assembly index does not imply minimum energy, and the assembly depth and the energy are independent complexity measures. Second, we ask how an ensemble T of strings co-assembles within a single space. Individuating step vertices by their formation histories rather than by the strings they carry yields three joint assembly spaces: the collectively (\( \omega_T \)), multiplicity (\( \mu_T \)), and singly (\( \pi_T \)) optimal ones, where the latter need not exist for certain ensembles and \( |\omega_T| \le |\mu_T| \le |\pi_T| \). All three sizes can differ, and multiplicity can lower the singly optimal size only for ensembles with \( |\pi_T| - |\omega_T| \ge 2 \). Collective optimization is NP-complete, while singly and multiplicity optimizations are NP-hard and lie a priori in the second level of the polynomial hierarchy.

Brief Report
Computer Science and Mathematics
Data Structures, Algorithms and Complexity

Pietro Hiram Guzzi

,

Marianna Milano

Abstract: Global network alignment maps nodes of one network onto nodes of another, preserving as much topology as possible. When two or more alignments are produced for the same pair of networks---by different algorithms or different random seeds of the same algorithm---a natural question arise: show much do these alignments agree? The standard answer, the Jaccard index over the sets of aligned node pairs, captures only set overlap and is blind to the local geometric structure that each alignment induces. We introduce a family of curvature-based agreement measures that compare alignments through the discrete graph curvatures they induce on the two networks. For each of three discrete Ricci-type curvatures—Forman–Ricci (kFR), Ollivier–Ricci (kOR), and Cayley–Menger (kCM)—we define two complementary agreement scores: a distribution-based \( \chi^2 \) distance and a correspondence-aware Spearman correlation. We evaluate these measures on 22 pairwise comparisons across 10 organism pairs (6 bacterial, 4 eukaryotic), using alignments produced by SANA and MAGNA++. We find that (i) no single curvature dominates across all comparisons (kCM is highest in 11/22, kFR in 8/22, kOR in 3/22); (ii) the Jaccard index replicates the finding of Guerra et al. (2020) that real alignments match the random null; (iii) the curvature-based Spearman measures detect agreement variation that Jaccard cannot, particularly in eukaryote pairs where the range widens to \( \rho \in [0.26, 0.93] \); and (iv) an approximate kCM based on truncated eigendecomposition scales to 13,276-node networks in \( \sim \)3.4 min while retaining Spearman \( \rho = 0.92 \) against the exact computation on a benchmark graph. These results establish discrete curvature as a viable and informative lens for alignment agreement, complementary to the established set-overlap paradigm.

of 13