Preprint
Article

This version is not peer-reviewed.

CAPPERI: Context-Aware and Human-Guided Composition of Data Preparation Pipelines

A peer-reviewed article of this preprint also exists.

Submitted:

07 September 2026

Posted:

08 September 2026

You are already at the latest version

Abstract
Data preparation requires heterogeneous operations to be composed into coherent pipelines while respecting application-specific constraints. Existing approaches span interactive curation, automated pipeline search, provenance, and operator recommendation, yet a complementary need remains for systems that preserve human control while enforcing validity during composition. CAPPERI (Context-Aware data Preparation Pipelines through Explicit Reuse and human Interaction) is an open-source software-intensive framework based on reusable abstractions for context, object type, algorithm, and pipeline. Data types and transformations form a typed directed graph, while explicit hierarchical application contexts determine the admissible portion of the graph and enable reuse through specialization. Step-wise guidance exposes only transformations compatible with the current data state and context, while a modular architecture separates domain abstractions, composition logic, validation, persistence, and interaction. Unlike learned-context automated search, CAPPERI treats context as explicit domain knowledge and an active composition constraint, combining automatic validity enforcement with human intentionality.
Keywords: 
;  ;  ;  ;  ;  ;  ;  

1. Introduction

Data preparation transforms raw, heterogeneous, incomplete, or noisy data into representations suitable for analysis. It is routinely recognized as one of the most labor-intensive parts of data science, and the challenge is not limited to implementing individual cleaning or transformation operators. A practitioner must decide which operations are appropriate, in which order they can be safely applied, and how generic preparation knowledge should be adapted to the application at hand. Contemporary reviews therefore describe data preparation as a pipeline-level problem spanning profiling, matching, mapping, transformation, repair, workflow construction, interaction, and partial automation [7].
Research has attacked different portions of this problem. Interactive systems such as Alpine Meadow support users during machine-learning pipeline curation and optimization [15]. Automated approaches search large spaces of preprocessing pipelines to maximize downstream predictive performance, including differentiable search [12], combinations of human- and machine-generated pipelines [8], and context-aware automated construction based on learned dataset representations [11]. Other work starts after a pipeline has been defined: fine-grained provenance can make preprocessing steps inspectable and help explain how data items were created, transformed, or removed [6], while provenance-based screening can identify correctness and compliance concerns in machine-learning preparation pipelines [14]. Recent work on scientific workflows has also explored next-operator suggestion from historical workflows [9]. At a larger semantic scale, pipeline metadata knowledge graphs have been proposed to support search and recommendation over large collections of AI pipelines [17]. A complementary body of work shows why domain meaning should enter data preparation and composition explicitly rather than remain implicit in operator names or dataset structure. [16] uses conceptual models to inject domain knowledge into machine-learning data preparation, with the aim of preserving domain semantics and improving both model performance and process transparency. In context-aware service composition, semantic context models have similarly been used to make context a first-class design concept that constrains discovery and composition [10]. More recently, contextualized ontology networks have been used to make entity context explicit, represent data semantics in context, and support knowledge reuse across different contexts [13]. These results motivate a distinction that is central to our contribution: context is valuable not only as information for ranking alternatives, but also as an explicit representation of the domain assumptions under which a transformation is meaningful and reusable. Earlier works on multimedia information retrieval provide a complementary precedent for this view. Context-based image similarity showed that the same data object may support different operational meanings depending on the surrounding context and on the user’s current task [2]. Subsequent work organized semantic annotations through multiple tree-structured taxonomies, using hierarchical concepts to represent progressively more specific interpretations of image content [3]. These results reinforce the idea that context and hierarchical conceptual organization can qualify meaning rather than merely annotate data after processing.
These advances leave room for a different design point. Fully automatic pipeline search is attractive when the objective can be expressed by a downstream metric and the transformation space is known in advance. In many data-management settings, however, the person constructing the pipeline has domain knowledge that should remain part of the decision process. At the same time, unrestricted manual composition is error-prone: an operator may be incompatible with the current representation, may belong to a different application domain, or may create a structurally invalid path. This motivates a human-guided approach in which the system does not choose the pipeline on behalf of the user, but continuously restricts the choice set to transformations that are admissible in the current state.
Pipeline composition is also a software-engineering problem. A practical environment must coordinate domain abstractions, transformation catalogs, validation rules, persistence, interactive state, and user-facing services without coupling these concerns into a monolithic implementation. Reuse is equally important: common processing knowledge should be defined once and specialized across related domains rather than repeatedly encoded in separate pipelines. Accordingly, a pipeline-composition environment can be viewed as a software-intensive system whose architecture should make the conceptual model, composition mechanism, and reusable abstractions explicit and independently extensible.
This paper presents CAPPERI (Context-Aware data Preparation Pipelines through Explicit Reuse and human Interaction),1 a context-aware framework for guided composition of reusable data-preparation pipelines. The design was developed in the S-PIC4CHU project, whose broader goal is to support semantics-aware, quality-oriented data preparation. The project view treats data preparation as a sequence of modular building blocks whose composition may require human intervention, and emphasizes that preparation should be tailored to application requirements rather than performed only at the syntactic level [1]. CAPPERI operationalizes a lightweight version of this idea: context is represented explicitly and actively constrains the transformations available during composition. In the current implementation, this representation is deliberately simple: a context is not an ontology and does not provide logical reasoning over domain concepts; rather, it captures a selected portion of application semantics operationally, by making explicit the scope in which object types and algorithms are meaningful and by organizing such scopes through specialization. This design occupies an intermediate point between purely syntactic type compatibility and fully ontology-based semantic composition.
Contributions of the paper are as follows:
1.
Reusable software abstractions based on Context, ObjectType, Algorithm, and Pipeline separate domain-level processing knowledge from a particular application implementation.
2.
Composition is formulated as navigation of a typed transformation graph, with step-wise validity checks preventing incompatible transformations and cycles.
3.
Explicit hierarchical application contexts act both as composition constraints and as a reuse-through-specialization mechanism: general elements are inherited by specialized domains, while domain-specific transformations remain locally scoped.
4.
These concepts are realized in a modular software architecture that separates orchestration, composition logic, validation, runtime models, persistence, catalog management, and user interaction.
The apparently closest recent work makes the design gap particularly clear. CtxPipe [11] addresses context-aware data-preparation pipeline construction through a fully automated paradigm: contextual information is extracted from the input data through pretrained embedding models and incorporated into reinforcement-learning-based pipeline search. CAPPERI uses the word context in a fundamentally different sense. Context is not a latent representation learned from a dataset; it is an explicit, first-class representation of an application domain. It determines which object types and algorithms are visible during composition, and its hierarchy specifies how preparation knowledge is inherited and reused across related domains. Consequently, context does not merely help rank candidate operators: it defines the admissible composition space itself.
This distinction also leads to a different allocation of responsibilities between automation and the user. Automated pipeline-search approaches ask the system to select a complete or high-performing preparation pipeline, typically against a downstream objective. CAPPERI deliberately retains the domain expert in the decision loop. Automation is used to enforce hard structural and contextual validity constraints, whereas the user retains intentional control over which admissible transformation should be applied. In short, automation handles validity, while humans retain intentionality. This design is especially relevant when appropriateness depends on domain goals that cannot be faithfully reduced to a single predictive-performance objective.
The contribution therefore does not rely on a claim of being the first “context-aware” pipeline system. Instead, CAPPERI operationalizes a different notion of context-aware composition: explicit hierarchical application context as an active constraint on human-guided pipeline construction. This position is complementary to learned-context automation such as CtxPipe and to next-operator recommendation systems such as FlowPilot [9].

2. Materials and Methods

2.1. Design Goals

CAPPERI was designed around four requirements.
Structural consistency 
requires that the output type of every step match the input type of the next transformation.
Guided composition 
requires that the user see only transformations that can be applied in the current state.
Context awareness 
requires that the available types and transformations depend on the active application context.
Reuse 
requires generic elements to be defined once and inherited by more specialized contexts.
The system intentionally follows a human-in-the-loop design. It does not attempt to infer a globally optimal pipeline or optimize downstream model accuracy. Instead, it constrains the feasible composition space and leaves the final choice among valid alternatives to the user. This choice is useful when pipeline appropriateness depends on domain goals that are difficult to reduce to a single objective function, and it also provides a clear separation between validity (which the system can enforce) and preference (which may remain with the domain expert).

2.2. Context-Aware Pipeline Model

The model contains four entities (Figure 1).
1.
An ObjectType denotes a data representation or state, such as RawData, MedicalImage, or TimeSeries.
2.
An Algorithm is a typed transformation a : s → t , where s and t are ObjectTypes, with an associated non-negative cost.
3.
A Pipeline is an ordered sequence of algorithms whose adjacent input and output types are equal.
4.
A Context is a logical application scope, each context bringing its own ObjectTypes and Algorithms; existing contexts form a rooted hierarchy.
For a context c, let A n c ( c ) denote c and all its ancestors. The visible ObjectTypes and Algorithms are the union of elements declared in contexts c ′ ∈ A n c ( c ) . Thus, a specialized context inherits generic preparation knowledge without copying it. Conversely, a transformation declared in a child context is not visible from its parent or from a sibling branch. From a semantic perspective, the hierarchy can be interpreted as a lightweight contextualization mechanism. An ancestor Context captures processing knowledge intended to remain meaningful across a broader domain, whereas a descendant narrows that domain and introduces additional types and transformations whose meaning is local to the specialization. This interpretation is consistent with work in which explicit context is used to qualify semantics and enable reuse across contexts [13], and with earlier multimedia models in which context disambiguates the semantic intent associated with the same query object [2]. The use of hierarchical concept taxonomies for semantic image annotation further provides a precedent for organizing meaning through progressively specialized conceptual structures [3]. CAPPERI remains intentionally less expressive than ontology-based context models used for semantic service composition [10].
The implementation represents ObjectTypes as immutable runtime objects and Algorithms as immutable typed edges. A Pipeline maintains its start type and ordered steps. Adding a step first retrieves the current output type and rejects an Algorithm whose input type differs. The pipeline cost is the sum of the costs associated with its algorithms2 and the current overall cost is displayed during composition. To achieve sufficient generality, ObjectTypes (and thus Algorithms and relative costs) can be parameterized, for example to model the length of a time series or the dimensionality of a vector.

2.3. Guided Composition Algorithm

A composition session starts from three user choices: active context c, start ObjectType s, and target ObjectType t, with the constraint that both endpoint types must be visible in the hierarchy of the active context, A n c ( c ) . The backend then obtains an Algorithm Catalog A c specific for context c and initializes a session containing the target t and the sequence P of selected algorithms.
At each step, the current pipeline is reconstructed from the session state. The catalog A c is queried by current input type c u r r e n t ( P ) , and the resulting set is returned to the user interface. If no admissible transformation remains before the target is reached, the session is reported as stuck. If the target ObjectType t is reached, the user can select to terminate the pipeline composition and let the sequence persist.
Algorithm 1 conceptually summarizes the overall procedure.
Algorithm 1 Guided composition procedure (conceptual pseudocode)
Input: context c, source type s, target type t
1. Build catalog A c containing algorithms visible from A n c ( c ) (i.e., c and its ancestors)
2. Initialize pipeline P = 〈 〉 ; c u r r e n t ( P ) = s ;
3. Repeat
4.      Build C = { a ∈ A c ∣ i n p u t ( a ) = c u r r e n t ( P ) }
5.      If C = ∅ , report a dead end; otherwise expose C to the user.
6.      Let the user select a ∈ C ; append a to P and let c u r r e n t ( P ) = o u t p u t ( a )
7. Until c u r r e n t ( P ) = t and the user is satisfied
8. Save the completed pipeline
This procedure makes human intervention explicit but bounded. A user cannot select an arbitrary operator from the global repository; choices are first restricted by context, then by input-type compatibility. In this sense, context is not metadata attached to a completed workflow; rather, it changes the set of actions available while the workflow is being constructed.

2.4. Algorithm Catalog

The Algorithm Catalog A c is the runtime structure used to retrieve applicable transformations. It stores the visible algorithms and maintains dictionary-based indexes by input and output ObjectType. In the current implementation, a lookup by input type therefore avoids scanning the entire catalog. The catalog is constructed from the persistent database after applying the context hierarchy and ObjectType filters, then cached using the context hierarchy as key. Cache invalidation is triggered when ObjectTypes or Algorithms change.
The indexed design matters because guided composition repeatedly asks the same logical question: “which transformations accept the current type?” The catalog moves this operation away from a global linear search by exploiting indexing at the underlying database level, thus allowing scaling when both querying and inserting ObjectTypes and Algorithms.

2.5. Software Architecture and Implementation

Figure 2 shows the implemented architecture. The backend is written in Python and exposes REST endpoints through FastAPI. SQLAlchemy maps persistent entities to SQLite tables. Pydantic schemas validate API requests and responses. Runtime classes implement the logical model independently of persistence. A Streamlit frontend manages the user session and invokes the backend without embedding the composition logic itself.
The Pipeline API maintains temporary composition sessions server-side. Its main operations initialize a session, return admissible next steps, accept a selected step, save a completed pipeline, and retrieve saved pipelines. The separation between runtime models and ORM entities allows the validity rules to be expressed independently of database representation. Database constraints additionally prevent duplicate ObjectTypes and Algorithms within a context and enforce non-negative costs.
The current release is a composition environment, not yet a pipeline execution engine. Algorithms are represented by name, typed input/output signatures, context, and cost; they are not yet bound to executable functions or external services. This boundary is intentional in the present study because it isolates the composition problem. Binding Algorithms to executable modules and adding an execution engine are natural extensions discussed later.

2.6. Implementation Availability

The prototype was developed and tested with Python 3.10.11, FastAPI 0.136.3, Streamlit 1.58.0, SQLAlchemy 2.0.50, Pydantic 2.13.4, and SQLite 3.40.1. The source repository contains setup scripts, dependency versions, API documentation support, example database initialization, and the complete backend and frontend implementation. The implementation is already maintained in a public GitHub repository at URI https://github.com/mpatella/CAPPERI.

3. Results and Discussion

3.1. Multi-Domain Case Study

The CAPPERI repository also includes an example database representing a case study of a simple healthcare setting (Figure 3). Hospital is the root context and provides reusable data-processing abstractions that are inherited by progressively specialized contexts. Diagnostics and Monitoring extend the root context, with the former modeling processing of, say, clinical images, while the latter is used for sensor-based analysis. CTscan and MRI specialize Diagnostics, while Sport specializes Monitoring. Generic types such as RawData, CleanData, FeatureData, and Report are declared at the root, whereas domain-specific types are introduced only where needed, such as MedicalImage under Diagnostics and GPSData under Sport.
The hierarchy illustrates the CAPPERI principle of reuse through specialization, whereby common abstractions are inherited and domain-specific elements remain scoped to the corresponding branch.
The seed configuration contains 19 ObjectTypes and 18 Algorithms globally. Context scoping substantially reduces the catalog exposed to a user: at the root, only three generic Algorithms are visible, an 83.3% reduction relative to the global catalog; first-level contexts expose six Algorithms (66.7% reduction), while second-level contexts expose nine (50% reduction). Importantly, reduction of the visible composition space induced by context does not remove generic transformations: reusable transformations defined in ancestor contexts remain available through inheritance.
The case study also demonstrates reuse at the pipeline level. A generic RawData -> CleanData -> FeatureData -> Report pipeline is defined in Hospital and remains meaningful to descendants. Diagnostics defines an image-oriented path from RawData to Diagnosis, while CTscan and MRI extend the inherited image representation with specialized reconstruction and analysis operators. Monitoring introduces a sensor-processing branch, and Sport reuses its inherited SensorData type before adding GPS and activity-specific steps.

3.2. Step-Wise Guidance

Figure 4 illustrates the composition interface of the CAPPERI pipeline builder. Once a context and endpoint types have been selected, the frontend requests the admissible transformations for the current state. The current ObjectType, accumulated cost, selected steps, and available Algorithms are displayed explicitly. After a selection, the state is updated and the process repeats until the target is reached or no valid continuation exists.
This interaction differs from both unrestricted workflow editors and fully automatic pipeline search. The system does not rank or silently select a next operation; it defines a valid local choice set. This makes the user’s decision traceable and leaves room for domain-specific preferences that are not yet encoded in the system.

3.3. Context as an Active Composition Constraint

The main design claim of CAPPERI is that context should influence pipeline construction before a pipeline exists. This differs from attaching domain labels or provenance metadata to a completed workflow. The active context determines which ObjectTypes and Algorithms can participate in the composition session, and hierarchical inheritance provides a controlled mechanism for reuse.

3.3.1. Different Notions of Context: CAPPERI Versus CtxPipe

The comparison with CtxPipe [11] is particularly informative because both systems use the term context-aware, but they operationalize context at different abstraction levels and for different purposes. Table 1 summarizes the distinction.
CtxPipe therefore treats context mainly as a learned signal for optimization; CAPPERI treats context as explicit knowledge organization and a hard composition constraint. The latter makes the reason for operator availability directly inspectable: an Algorithm is visible because it belongs to the active Context or an ancestor and because its input type matches the current data state. This provides interpretability and controllability at composition time. The trade-off is that domain contexts must currently be curated explicitly rather than inferred automatically.
This distinction is not merely terminological. A transformation may be syntactically applicable to a representation while still being inappropriate for a particular application domain. By scoping types and transformations before composition, CAPPERI can represent such domain boundaries independently of any downstream ML score. Conversely, learned contextual representations could eventually enrich, but need not replace, this explicit domain model.
The same broader distinction applies to HAIPipe and DiffPrep. HAIPipe combines human-generated and machine-generated preparation pipelines [8], while DiffPrep searches differentiable preprocessing spaces for tabular learning [12]. Those systems primarily aim to discover high-performing pipelines for downstream ML. CAPPERI instead aims to make heterogeneous preparation operators safely composable across application domains, including settings where there may be no downstream ML objective at composition time.

3.4. Human Guidance Rather than Full Automation

Interactive curation has long been motivated by the mismatch between statistical/ML expertise and domain expertise [15]. CAPPERI takes a deliberately lightweight position: rather than automatically generating a pipeline and asking the user to validate it, the system narrows the decision space and lets the user choose among valid transformations. This can be understood as a form of constrained human agency.
FlowPilot [9] represents a recent and complementary direction: it recommends the next operator in scientific workflow development using historical workflow knowledge and user feedback. The three paradigms can be distinguished by the question they answer:
CtxPipe 
asks which preparation pipeline should be generated automatically for the dataset and downstream task;
FlowPilot 
asks which operator is likely to be useful next;
CAPPERI 
asks which transformations are admissible here, given the explicit application context and current data state.
This separation suggests a natural two-stage architecture for future assistance. CAPPERI can first apply deterministic context and type constraints to establish the admissible set; a recommender can then rank only those valid alternatives using historical pipelines, semantic similarity, cost, or user preferences. Recommendation and constrained composition are therefore orthogonal: soft ranking can improve guidance without being allowed to violate hard validity constraints.

3.5. Relation to Provenance and Semantic Pipeline Descriptions

Fine-grained provenance work shows why preprocessing pipelines must be treated as first-class objects: preparation steps affect not only accuracy and performance but potentially fairness and other downstream properties [6]. Provenance-based screening extends this perspective toward automated checks over preparation pipelines [14]. CAPPERI addresses the upstream side of the lifecycle. It constrains which transformations can be selected during design; provenance systems can subsequently record what those transformations did during execution.
The semantic direction is equally relevant. The literature provides evidence at several abstraction levels. Conceptual modeling can carry domain knowledge into data preparation so that preparation decisions better preserve what the data represents [16]. Semantic context models can make context a first-class constraint during service discovery and composition [10], while contextualized ontology networks can represent data semantics relative to explicit contexts and promote ontology reuse across them [13]. Together, these works indicate that context and semantics are related but not interchangeable: semantics describes the meaning of data and operations, whereas context qualifies the circumstances or domain scope under which that meaning is intended to hold. The same distinction has also emerged in multimedia data management. Context-based retrieval demonstrated that interpretation can depend on the user’s current semantic intent [2], while semantic taxonomies can bridge low-level representations and conceptual descriptions [3]. More broadly, combining visual content with explicit semantics has been advocated as a way to overcome the semantic gap that arises when automatically extracted features alone cannot capture the meaning relevant to users [4]. These observations are consistent with the role assigned to Context in CAPPERI: the framework does not infer full domain semantics, but makes the application scope in which types and transformations are intended to be meaningful explicit and operational. [17] demonstrates that explicit pipeline metadata and semantic enrichment can support search and recommendation over very large pipeline collections. Our Context abstraction is intentionally lighter than an ontology or knowledge graph, but the separation between context, types, algorithms, and pipelines creates natural attachment points for richer semantics. Accordingly, the claim made here is not that CAPPERI already performs semantic reasoning: its contribution is to make application scope explicit and operational at composition time, thereby providing a software abstraction in which richer semantic descriptions can subsequently be attached without changing the basic human-guided composition paradigm. In the S-PIC4CHU vision, where CAPPERI originates from, semantic information should accompany multiple stages of a Data Preparation Pipeline and support coherent, application-aware preparation [1]. A future version of CAPPERI can therefore replace or enrich string-level ObjectTypes with ontology-backed concepts and semantic compatibility relations.

3.6. Multimedia and Stream-Processing Perspective

The framework was motivated in part by data-preparation settings where representations evolve through multiple stages, including multimedia. Previous work on multimedia retrieval and stream processing has emphasized both feature-level representations and abstraction over distributed processing infrastructures. In particular, SPAF [5] provides an abstraction layer for constructing real-time stream-processing applications independently of specific stream-processing engines. CAPPERI operates at a different level: it focuses on selecting and composing typed preparation steps rather than executing a streaming topology. The two abstractions are compatible: binding CAPPERI Algorithms to SPAF processors or other executable operators would turn a validated logical preparation path into an executable pipeline while preserving separation between design and runtime infrastructure.

3.7. Software Architecture, Reusable Abstractions, and Extensibility

The contribution of CAPPERI is not tied to the specific technologies used by the current prototype. FastAPI, SQLAlchemy/SQLite, Pydantic, and Streamlit provide a concrete realization, but the reusable core lies in the separation between domain abstractions, composition rules, catalog services, persistence, and interaction. This separation of concerns is important for a software-intensive system because each layer can evolve without redefining the conceptual model of pipeline composition.
The four abstractions also support reuse at two levels:
At the model level, 
ObjectType and Algorithm describe processing knowledge independently of a particular pipeline instance.
At the domain level, 
hierarchical Contexts implement reuse through specialization: generic types and transformations are defined once in an ancestor and automatically become available in descendants, whereas specialized contexts add only the knowledge peculiar to their domain. The healthcare hierarchy used in the evaluation illustrates this mechanism through shared hospital-level operations and progressively specialized diagnostic, imaging, monitoring, and sport-processing branches.
From a software-engineering perspective, this design creates a number of explicit extension points:
  • new persistence technologies can replace SQLite without changing runtime composition rules;
  • alternative user interfaces can consume the REST API;
  • richer type systems can extend ObjectType compatibility;
  • executable services can be bound to Algorithm descriptors;
  • recommendation or provenance components can be layered around the existing deterministic validity mechanism.
Reusable abstractions and modular architecture therefore serve not only maintainability but also the controlled evolution of the system toward richer semantics and execution capabilities.

3.8. Opportunities for Improvement

We envisage to improve the current prototype along four different dimensions.
1.
Algorithms are descriptors rather than executable operators, so the system validates composition but does not yet execute the resulting pipeline.
2.
ObjectType compatibility is exact: semantic subsumption, coercion, parameter constraints, and multi-input/multi-output transformations are not currently represented.
3.
Cost is additive and informational only; no automatic cost-based path optimization is implemented.
4.
Contexts are manually curated and do not yet exploit ontologies, embeddings, or learned metadata.
The evaluation is correspondingly focused on the implemented contribution: context-based search-space restriction, guided validity, reuse in the supplied hierarchy, and catalog retrieval behavior. It does not claim improved downstream model accuracy because the framework does not execute ML pipelines. This distinction is important when comparing against automated pipeline-search systems whose primary outcome is predictive performance.

3.9. Future Development

The immediate engineering extension is to associate each Algorithm descriptor with an executable function, script, container, or service and introduce an execution engine that passes the output of one step to the next. A second extension is richer semantic typing, where compatibility can be inferred through ontology relations rather than exact ObjectType equality. Third, the existing cost attribute can support multi-criteria recommendation or optimization once costs acquire explicit semantics such as latency, monetary cost, energy, or quality impact. Fourth, provenance capture can be integrated with execution so that a saved pipeline records not only its design but also the lineage of produced data. Finally, recommendation can be layered on top of deterministic context/type filtering, learning from previously saved pipelines without allowing the recommender to violate hard composition constraints.

4. Conclusions

CAPPERI addresses a specific gap between unrestricted manual workflow design and fully automated pipeline search. It models data-preparation composition as navigation through a context-scoped typed transformation graph, keeps the domain expert in the loop, prevents structurally invalid steps, and supports reuse through hierarchical contexts. The implementation demonstrates that these abstractions can be realized in a modular software architecture and that context scoping can substantially reduce the transformation catalog presented to users. Indexed lookup further supports interactive retrieval of applicable operators. The broader implication is that context can be treated as an operational part of pipeline composition rather than passive documentation. This provides a foundation on which richer semantics, recommendation, provenance, execution, and quality-aware optimization can be added incrementally. By separating hard validity constraints from human choice and future soft recommendations, the framework offers a transparent path toward more intelligent yet controllable data-preparation environments.

Author Contributions

Conceptualization, I.B. and M.P.; methodology, I.B. and M.P.; validation, I.B. and M.P.; formal analysis, I.B. and M.P.; investigation, I.B. and M.P.; resources, I.B. and M.P.; writing—original draft preparation, I.B. and M.P.; writing—review and editing, I.B. and M.P.; visualization, I.B. and M.P.; supervision, I.B. and M.P.; project administration, I.B. and M.P.; funding acquisition, I.B. and M.P. All authors have read and agreed to the published version of the manuscript.

Funding

This research was funded by the Italian Ministry of University and Research (MUR) under PRIN 2022 grant 2022XERWK9, “S-PIC4CHU—Semantics-based Provenance, Integrity, and Curation for Consistent, High-quality, and Unbiased data science.”

Institutional Review Board Statement

Not applicable.

Data Availability Statement

No new research dataset was generated for this software study. The source code and example configuration used to reproduce the reported software behavior are maintained in a public GitHub repository at URI https://github.com/mpatella/CAPPERI.

Acknowledgments

The student Samuele Mazziotti is gratefully acknowledged for his contribution to the software implementation of the CAPPERI prototype. The S-PIC4CHU project partners are also acknowledged for providing interesting discussions on semantics-aware data preparation. During the preparation of this manuscript, the authors used OpenAI ChatGPT (GPT-5.6 Sol) for grammar and spelling review. After using this service, they carefully reviewed and edited the manuscript as needed and take full responsibility for the content of this publication.

Conflicts of Interest

The authors declare no conflict of interest. The funders had no role in the design of the study; in the collection, analyses, or interpretation of results; in the writing of the manuscript; or in the decision to publish the results.

References

  1. Alfano, G.; Bartolini, I.; Calvanese, D.; Ciaccia, P.; Greco, S.; Lanti, D.; et al. S-PIC4CHU: Semantics-Based Provenance, Integrity, and Curation for Consistent, High-Quality, and Unbiased Data Science. In Proceedings of the 33rd Symposium on Advanced Database Systems (SEBD), 2025. [Google Scholar]
  2. Bartolini, I. Context-Based Image Similarity Queries. In Adaptive Multimedia Retrieval: User, Context, and Feedback;Lecture Notes in Computer Science; Springer: Berlin/Heidelberg, Germany, 2006; Volume 3877, pp. 222–235. [Google Scholar] [CrossRef]
  3. Bartolini, I.; Ciaccia, P. Multi-Dimensional Keyword-Based Image Annotation and Search. In Proceedings of the 2nd International Workshop on Keyword Search on Structured Data (KEYS 2010), Indianapolis, IN, USA, 6 June 2010; pp. 5:1–5:6. [Google Scholar] [CrossRef]
  4. Bartolini, I. Content Meets Semantics: Smarter Exploration of Image Collections—Presentation of Relevant Use Cases. In Proceedings of the International Conference on Signal Processing and Multimedia Applications (SIGMAP 2012), Rome, Italy, 24–27 July 2012; pp. 186–191. [Google Scholar] [CrossRef]
  5. Bartolini, I.; Patella, M. A Stream Processing Abstraction Framework. Front. Big Data 2023, 6, 1227156. [Google Scholar] [CrossRef]
  6. Chapman, A.; Missier, P.; Simonelli, G.; Torlone, R. Capturing and Querying Fine-Grained Provenance of Preprocessing Pipelines in Data Science. Proc. VLDB Endow. 2021, 14, 507–520. [Google Scholar] [CrossRef]
  7. Chapman, A.; Simperl, E.; Koesten, L.; Konstantinidis, G.; Ibanez, L.D.; Kacprzak, E.; et al. Data Preparation: A Technological Perspective and Review. SN Comput. Sci. 2023, 4. [Google Scholar] [CrossRef]
  8. Chen, S.; Tang, N.; Fan, J.; Yan, X.; Chai, C.; Li, G.; Du, X. HAIPipe: Combining Human-Generated and Machine-Generated Pipelines for Data Preparation. Proc. ACM Manag. Data 2023, 1, 91:1–91:26. [Google Scholar] [CrossRef]
  9. Esmailoghli, M.; Weidlich, M. FlowPilot: A Suggestion System for Designing Scientific Workflows. Proc. ACM Manag. Data 2026, 4, 39. [Google Scholar] [CrossRef]
  10. Furno, A.; Zimeo, E. Context-Aware Composition of Semantic Web Services. Mob. Netw. Appl. 2014, 19, 235–248. [Google Scholar] [CrossRef]
  11. Gao, H.; Cai, S.; Dinh, T.T.A.; Huang, Z.; Ooi, B.C. CtxPipe: Context-Aware Data Preparation Pipeline Construction for Machine Learning. Proc. ACM Manag. Data 2024, 2, 231:1–231:27. [Google Scholar] [CrossRef]
  12. Li, P.; Chen, Z.; Chu, X.; Rong, K. DiffPrep: Differentiable Data Preprocessing Pipeline Search for Learning over Tabular Data. Proc. ACM Manag. Data 2023, 1, 183:1–183:26. [Google Scholar] [CrossRef]
  13. Rico, M.; Taverna, M.L.; Galli, M.R.; Caliusco, M.L. Context-Aware Representation of Digital Twins’ Data: The Ontology Network Role. Comput. Ind. 2023, 146, 103856. [Google Scholar] [CrossRef]
  14. Schelter, S.; Guha, S.; Grafberger, S. Automated Provenance-Based Screening of ML Data Preparation Pipelines. Datenbank-Spektrum 2024, 24, 187–196. [Google Scholar] [CrossRef]
  15. Shang, Z.; Zgraggen, E.; Buratti, B.; Kossmann, F.; Eichmann, P.; Chung, Y.; et al. Democratizing Data Science through Interactive Curation of ML Pipelines. Proc. ACM Manag. Data, 2019; pp. 1171–1188. [Google Scholar] [CrossRef]
  16. Storey, V.C.; Parsons, J.; Castellanos Bueso, A.; Chiarini Tremblay, M.; Lukyanenko, R.; Castillo, A.; Maass, W. Domain Knowledge in Artificial Intelligence: Using Conceptual Modeling to Increase Machine Learning Accuracy and Explainability. Data Knowl. Eng. 2025, 160, 102482. [Google Scholar] [CrossRef]
  17. Venkataramanan, R.; Tripathy, A.; Kumar, T.; Serebryakov, S.; Justine, A.; Shah, A.; et al. Constructing a Metadata Knowledge Graph as an Atlas for Demystifying AI Pipeline Optimization. Front. Big Data 2025, 7, 1476506. [Google Scholar] [CrossRef]
Figure 1. Conceptual model of CAPPERI: ObjectTypes are nodes of a typed transformation graph, Algorithms are directed edges, and a Pipeline is a valid path. A hierarchical Context determines which nodes and edges are visible during composition.
Figure 1. Conceptual model of CAPPERI: ObjectTypes are nodes of a typed transformation graph, Algorithms are directed edges, and a Pipeline is a valid path. A hierarchical Context determines which nodes and edges are visible during composition.
Preprints 232101 g001
Figure 2. CAPPERI software architecture: the user interface is separated from the orchestration, validation, runtime model, catalog, and persistence layers.
Figure 2. CAPPERI software architecture: the user interface is separated from the orchestration, validation, runtime model, catalog, and persistence layers.
Preprints 232101 g002
Figure 3. Context hierarchy used in the supplied case study: elements defined in an ancestor context are inherited by descendants, while specialized elements remain scoped to their branch.
Figure 3. Context hierarchy used in the supplied case study: elements defined in an ancestor context are inherited by descendants, while specialized elements remain scoped to their branch.
Preprints 232101 g003
Figure 4. Example of the guided composition interface: the user is presented with the current state and only transformations returned as admissible by the backend.
Figure 4. Example of the guided composition interface: the user is presented with the current state and only transformations returned as admissible by the backend.
Preprints 232101 g004
Table 1. Different notions of context-aware pipeline construction in CtxPipe and CAPPERI.
Table 1. Different notions of context-aware pipeline construction in CtxPipe and CAPPERI.
Dimension CtxPipe CAPPERI
Meaning of context Learned information extracted from the input dataset Explicit application/domain scope represented as a first-class entity
Representation Pretrained embedding-based contextual representation Named hierarchical Contexts containing/inheriting ObjectTypes and Algorithms
Role in composition Informs automated search and component selection Determines the admissible transformation space before user selection
Composition paradigm Fully automated pipeline construction Human-guided, step-wise constrained composition
Decision mechanism Reinforcement-learning-based search toward downstream performance Deterministic context filtering and typed compatibility, followed by user choice
Primary objective Discover high-quality ML data-preparation pipelines Construct valid, interpretable, reusable data-processing/preparation pipelines
Reuse Learned search experience across datasets Explicit inheritance of preparation knowledge across related application contexts
Transparency/control Selection mediated by learned search policy Every admissible next transformation is exposed explicitly; the user retains final control
1
Besides its literal meaning (capers), in Italian “capperi” is also used as a popular, polite exclamation to express mild astonishment or admiration, e.g., “Capperi, che bell’articolo!” (Wow, what a nice paper!).
2
Pipeline cost is only informative and is not (yet) used for automatic optimization.
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.