Preprint
Article

This version is not peer-reviewed.

The Karaka Calling Convention: Semantic-Role Argument Binding for Logic Programming

Submitted:

23 August 2026

Posted:

25 August 2026

You are already at the latest version

Abstract
Positional argument binding, in which f(x, y, z) assigns meaning by slot order, degrades as arity grows, and logic programming suffers it worst: each position's semantics must be remembered rather than read off the call site, and keyword arguments only trade this for a vocabulary that is local to each predicate. We present the Karaka Calling Convention (KCC): arguments are bound by a fixed vocabulary of six semantic roles (agent, object, instrument, recipient, source, locus), drawn from Paninian Sanskrit grammar and carried morphologically by each argument. Because the vocabulary is predicate-independent, a rewrite rule written once against a role applies across all predicates, a property that keyword arguments and PropBank framesets lack. A parsed statement is already a labeled-edge graph, so the same object compiles to Prolog facts and to property-graph writes with no lowering pass, and groundness alone decides whether a statement is an assertion or a query. Rule conflicts resolve by requirement-set subsumption (the Paninian utsarga/apavada discipline), with incomparable rules rejected at registration as ambiguous. We give a formal definition, a worked end-to-end example, and a tested reference implementation with optional real Astadhyayi morphology via vidyut-prakriya; a prior-art search found one contemporaneous system using Karaka roles as call-site labels, and we characterize the overlap from its source. Measured readability benefit is open work, stated as such.
Keywords: 
;  ;  ;  ;  ;  ;  ;  

1. Introduction

The idea that Sanskrit's grammar has something to offer computer science is not new. Briggs [1] argued in AI Magazine that Paninian Karaka analysis, a system of six semantic case roles attached to nouns via case endings, produces structures functionally equivalent to the semantic nets then used in AI knowledge representation, and that because Sanskrit word order carries little more than stylistic weight, Karaka-marked sentences are naturally suited to a syntax-free representation: a flat list of role-tagged semantic relations.
Most subsequent "Sanskrit programming language" projects have taken a different and much shallower approach: translating a conventional language's keywords into transliterated Sanskrit while leaving positional, C-like syntax untouched. This addresses none of the structural claims in Briggs's argument and is, in effect, orthographic rather than architectural.
This paper takes the architectural claim seriously and narrowly: Karaka roles as a calling convention for logic/relational programming, where the positional-argument problem is sharpest. Panini supplies the role vocabulary and the rule-ordering discipline; he is not a source of validity for the claim itself. The observation everything else in this paper follows from is structural: a statement whose arguments carry their own roles is already a labeled-edge graph, and a representation that is already a graph needs no lowering pass to become either a fact base or a query.

3. Design: The Karaka Calling Convention

This section defines the convention itself: what a statement looks like, how an argument acquires its role, how the six classical roles extend to relations that need more, and why the resulting structure serves as a fact base and as a query without an intervening lowering pass. Sanskrit terms appear in plain ASCII transliteration throughout, matching the identifiers used in the reference implementation.
A KCC statement consists of one verb (predicate) and up to six role-marked arguments, karta (agent), karma (object), karana (instrument), sampradana (recipient), apadana (source), adhikarana (locus), plus an overflow sambandha (genitive/possessive) role for nested sub-frames. Each argument's role is determined by a marker carried on the token itself (in the reference implementation, a suffix), not by its position in the statement. The parse target is a Frame: a verb plus a {role: argument} mapping. Figure 1 shows the full pipeline, from surface statement through role resolution and rule application to compiled output. Three permutations of the same statement's word order produce an identical Frame by construction (verified by test, Section V).
Arity beyond the six classical roles does not require adding more roles: any role's value can itself be a nested Frame rather than a leaf string, so a single "instrument" can expand into an entire sub-event with its own roles (e.g. an instrument that is itself caused by a prior event). This generalizes cleanly to arbitrary nesting depth, though it does mean deeply nested KCC statements trade the flat readability of a two- or three-role frame for the structure of a small tree, a real trade, not a free win.
A Frame is, incidentally, already a small labeled-edge graph structure: an event node with role-labeled edges to its argument nodes, as in Figure 2. This motivates the two compilation targets evaluated here and shown side by side in Figure 3: normalized Prolog facts (one binary predicate per role, avoiding positional-term degradation) and Cypher MERGE statements against a property graph (e.g. FalkorDB). Why this matters, rather than being a curiosity: it collapses the usual distance between a program and a knowledge base. In a conventional stack, getting from source code to something a graph database or reasoner can consume requires an extraction pipeline, parse, walk the AST, decide on a schema, emit triples. Here there is nothing to extract: the parsed statement is the graph, so a KCC program is simultaneously executable (as facts and queries), storable (as a property graph), and pattern-matchable (by any graph or Datalog engine) with no representation change at any step. Both codegen paths recurse through nested frames, emitting a fresh sub-event identifier and linking it from the parent by the role name, a Skolemization step, in the standard logic-programming sense: an existentially-quantified sub-event ("there exists some flooding event that is the instrument here") is witnessed by a fresh constant, which is exactly what turns a nested, tree-shaped Frame into flat, indexable relational tuples without losing the parent-child relationship. We refer to this labeled-edge Frame representation as a semantic graph throughout.
A. Why this is not just keyword arguments with extra steps
Keyword arguments already give order-independence and readable call sites. The difference is a single property, which we state and then defend:
Proposition (predicate-independence).A role vocabulary is predicate-independent iff the interpretation of a role does not depend on which predicate it is attached to. A rewrite rule is predicate-agnostic iff its applicability is decidable from facts(F) alone, with no per-predicate lookup. A fixed, predicate-independent vocabulary is what makes predicate-agnostic rules well-defined.
Python keyword names, Rust named fields, and PropBank's :ARG0–:ARG5 (the largest deployed role-labeling system, underlying AMR) are all predicate-relative: PropBank's own framesets gloss :ARG0 as "wanter" under want-01, "sleeper" under sleep-01, and "creator" under develop-02, the same label, three meanings, resolved only through a per-verb table. A Karaka role carries no such indirection: karta means agent for every predicate, by construction.
The consequence is demonstrable, not aesthetic. The following rule names no verb:
SubsumptionSutra(
  name="mediated-transfer-classification",
  requires=frozenset({"has:karana",
     "has:sampradana"}),
  transform=lambda f: Frame(f.verb, {**f.roles,
     "class": "mediated-transfer"}))
Applied to two frames under different verbs, hara (causes harm) and dada (gives), it classifies both, and leaves a two-role motion frame untouched (output generated by the test suite, test_predicate_agnostic_rule_applies_across_different_verbs):
Frame(hara, class=mediated-transfer, karana=jal,
     karta=man, sampradana=ksetr)
Frame(dada, class=mediated-transfer, karana=dhan,
     karta=nar, sampradana=putr)
Written against keyword arguments or ARG-numbered roles, this requires either one clause per predicate or a maintained mapping from each local vocabulary to a shared one, and that mapping is a Karaka layer, reinvented.
Two honest qualifications. First, the vocabulary property alone is replicable anywhere: globally reserving agent, instrument, recipient as keyword names in any language reproduces it, and Prakash's Karaka (Section II) does essentially this with the Sanskrit terms. The vocabulary is therefore necessary for this paper's contribution but not sufficient; the contribution is the composition (Section VIII). Second, real Sanskrit complicates the surface mapping: passivization and idiosyncratic case government mean the vibhakti-to-Karaka correspondence is itself verb-conditioned in general (Section VI). That table, if added, is structurally a frameset, but it sits at the surface layer only. Its output is still the fixed six-role vocabulary, so every rule downstream of parsing remains predicate-agnostic; what becomes verb-conditioned is the mapping from case ending to role, not the meaning of the role.
B. Formal definition
The core structure is small enough to state completely:
Frame  ::= (Verb, Bindings)
Bindings ::= Role ⇀ Value       (a partial map;
     each role at most once)
Role   ∈  {karta, karma, karana, sampradana,
     apadana, adhikarana, sambandha}
Value  ::= Atom | Var | Frame
Var    ::= "?" Atom
A frame is ground if no Var occurs in it (recursively); ground frames denote assertions, non-ground frames denote queries (Section III-C). The fact set of a frame, used by the rule engine (Section IV), is defined recursively:
facts(F) = {verb=v} ∪ {has:r | r ∈ dom(F.bindings)}
       ∪ { r.φ | r ↦ F' ∈ F.bindings, F' a
     Frame, φ ∈ facts(F') }
Here verb=v records the frame's predicate, has:r records that role r is bound, and r.φ prefixes every fact φ of a nested sub-frame with the role it hangs under. So for a harm event whose instrument is itself a rainfall event, facts contains verb=hara, has:karta, has:karana, and the nested karana.verb=varsa, karana.has:karta, which is what lets a rule condition on "the instrument was a natural event." A rule's condition language is deliberately minimal: requires is a conjunction of these atomic facts, with no negation or disjunction; subsumption between rules is then plain set inclusion, which is what keeps the ordering of Section IV decidable at rule-registration time.
Theorem 1 (Order Invariance). 
Any permutation of the tokens of a well-formed KCC statement parses to an identical Frame.
Proof sketch. Each token carries its role marker morphologically, so role binding is a function of the token alone; the parser reads tokens as a set, rejecting duplicate roles and verb counts other than one, and the resulting Frame is a map from roles to values, which is order-free by construction. No parsing step consults position, so permutation cannot change the result. The property is exercised directly by the order-invariance tests in the reference implementation.
C. A worked example, end to end
The preceding subsections define the convention piecewise; this one runs a single statement through every stage of Fig. 1. The statement is a Rylands v. Fletcher-shaped tort relation, "the man causes harm to the field by means of the flood, from the reservoir":
manah ksetraya jalena jalasayat harati
Each token carries its role as a case suffix (-ah agent, -aya recipient, -ena instrument, -at source; -ti marks the verb), so the parser needs no positional information and any permutation of the five words parses identically:
>>> from karaka_lang import parse
>>> parse("manah ksetraya jalena jalasayat harati")
Frame(hara, apadana=jalasay, karana=jal, karta=man,
     sampradana=ksetr)
Two rules are then applied. In the implementation, a rule is a declared requirement set plus a transform:
SubsumptionSutra(
  name="default-liability",             #
     utsarga: the general rule
  requires=frozenset({"verb=hara"}),
  transform=lambda f: Frame(f.verb, {**f.roles,
     "liability": "strict"}))
SubsumptionSutra(
  name="instrument-liability-exception",     #
     apavada: the exception
  requires=frozenset({"verb=hara", "has:karana"}),
  transform=lambda f: Frame(f.verb, {**f.roles,
     "liability": "instrument-based"}))
The second rule's requirements strictly include the first's, so it is more specific and applies last, overriding on the shared output key. Had the two requirement sets been incomparable, the engine would raise AmbiguityError instead of choosing. The resolved frame then compiles, without an intermediate lowering pass, to both targets:
% to_prolog(frame): one binary predicate per role,
     arity never grows
event(e1, hara).
apadana(e1, jalasay).
karana(e1, jal).
karta(e1, man).
liability(e1, instrument-based).
sampradana(e1, ksetr).
// to_cypher(frame): the same shape as labeled edges
MERGE (e:Event {verb: "hara"})
MERGE (n_jalasay:Entity {name: "jalasay"}) MERGE
     (e)-[:APADANA]->(n_jalasay)
MERGE (n_jal:Entity {name: "jal"})      MERGE
     (e)-[:KARANA]->(n_jal)
MERGE (n_man:Entity {name: "man"})      MERGE
     (e)-[:KARTA]->(n_man)
MERGE (n_ksetr:Entity {name: "ksetr"})   MERGE
     (e)-[:SAMPRADANA]->(n_ksetr)
Finally, replacing any constant with a variable turns the same structure into a question. "Who causes harm by means of the flood?" is the same frame with karta left open, and it compiles to a query rather than a write:
>>> compile_frame(parse("?kah jalena harati"),
     target="prolog")
event(E, hara), karana(E, jal), karta(E, K).
>>> compile_frame(parse("?kah jalena harati"),
     target="cypher")
MATCH (e:Event {verb: "hara"})
MATCH (e)-[:KARANA]->(:Entity {name: "jal"})
MATCH (e)-[:KARTA]->(k:Entity)
RETURN k.name
Assertion and query are one type; groundness decides which you get. All output above is generated by the reference implementation, not typeset by hand.

4. Rule Resolution: From Manual Priority to Subsumption-Derived Order

A fixed role vocabulary makes predicate-agnostic rules expressible, but expressibility is not enough: when several rules match the same frame, something must decide the order in which they apply, and that decision determines whether independently written rule sets can be combined at all. This section describes how the order is derived from the rules themselves, and why an earlier design that assigned it by hand was replaced.
Panini's grammar is organized around general rules (utsarga) that can be overridden by more specific exception rules (apavada), the same principle underlying CSS specificity or pattern-match ordering in functional languages. An initial implementation used an explicit, author-assigned integer priority per rule. That design has a known weakness: manual integers do not compose. If two independently authored rule sets each assign priority=5, combining them produces an arbitrary ordering rather than a meaningful one.
The current implementation instead derives specificity from condition subsumption. Each rule declares a set of required conditions on the frame (e.g. {verb=hara, has:karana}); rule A is apavada to rule B iff A's requirements are a strict superset of B's, and applies later, overriding B on any conflicting output. Requirement-set inclusion induces a partial order over applicable rules. If two applicable rules are incomparable under this order, neither a subset of the other, the engine raises a structural AmbiguityError rather than guessing, mirroring the way Paninian commentators required an explicit tie-breaking rule when a conflict didn't resolve by the grammar's own ordering principles. This is deliberately stricter than common practice, many rule systems default to "last defined wins" or fall back to declaration order, and a configurable tie-breaker would be easy to add; we chose rejection as the default because silent, order-dependent resolution is precisely the composability failure the manual-priority version suffered from. This closes the composability gap the reviewers identified, though it is worth noting plainly that the "requirements" a rule declares are still author-written, not automatically extracted from the condition predicate itself, full automatic subsumption analysis over arbitrary Python predicates is not attempted, and would be a substantially larger undertaking. Equivalently, an AmbiguityError signals that the set of applicable rules has no unique maximum under this order.

5. Implementation and Evaluation

The design above is realized in a working implementation, which produced every code output shown in this paper. We describe what it contains, what the test suite establishes, and what it does not.
The reference implementation, test suite, examples, and all output reproduced in this paper are publicly available at https://github.com/joyboseroy/karaka-lang.
The reference implementation (karaka_lang, Python) consists of a tokenizer/role-resolver, the Frame AST, two rewrite engines (Section IV), codegen to Prolog and Cypher in both assertion and query mode (Section III-C), and an optional morphology module binding to vidyut-prakriya that replaces the toy suffix table with actual Astadhyayi derivation over a registered lexicon, every analyzed form carries the sutra numbers that derived it (e.g. aSvena resolves to aSva + instrumental → karana via the chain 1.2.45 → 4.1.2 → ... → 8.4.68). A 23-test suite covers: order invariance across permutations; duplicate-role and missing-verb error handling; the five-role tort relation of Section III-C; utsarga/apavada override under both engines, with and without a matching exception; rejection of incomparable rules via AmbiguityError; a predicate-agnostic rule applying across different verbs (Section III-A); rules conditioning on nested sub-event facts; nested-Frame arity extension with recursive codegen; query-mode compilation to both targets; and, when vidyut is installed, real-declension round-trips including rejection of out-of-lexicon and ambiguous forms. All 23 tests pass (20 without vidyut; the morphology tests skip cleanly). This remains a proof-of-concept evaluation, not a benchmark: no comparison against Prolog predicate readability or programmer performance has been conducted (Section VI).

6. Limitations

Several parts of this design are complete only in a restricted form, and the restrictions are load-bearing for how much the paper can claim. We therefore list them explicitly rather than in passing. Three are partial, in that a mechanism exists but covers a narrower case than the general statement of the problem; three are absolute, in that the work does not address them at all.
  • • Morphology, partially addressed. The optional vidyut-prakriya binding (Section V) performs real Astadhyayi derivation, but over a closed, registered lexicon (analysis-by-generation), singular forms only, and with the default vibhakti-to-Karaka correspondence. Real Sanskrit breaks that neat mapping under passivization and verbs governing idiosyncratic cases; a per-verb mapping table is the known fix and is future work. The zero-dependency toy suffix table remains the fallback.
  • • Whitespace tokenization. Real Sanskrit sandhi merges word boundaries; the reference implementation sidesteps this entirely. vidyut-cheda or the sanskrit_parser system [14]'s DP sandhi-split search are the correct components to integrate, and neither is trivial (both report ongoing over/under-generation issues). This remains fully open.
  • • Arity ceiling, addressed with a caveat. Nested Frames under a role remove the hard six-role ceiling (Section III), and the subsumption engine computes role-prefixed facts recursively, so a rule can condition on a property of a nested sub-event. What remains restricted is depth-independent quantification: a rule can require karana.verb=varsa at a known path, but cannot yet express "some sub-event at any depth has property P." Deeply nested statements also trade the flat readability of a two- or three-role frame for the structure of a small tree.
  • • Rule-conflict resolution, partially addressed. Specificity is now derived from declared condition sets rather than an author-assigned integer (Section IV), and genuinely incomparable rules raise AmbiguityError rather than resolving silently. It does not, however, extract those condition sets automatically from arbitrary rule logic. A rule's requires is still hand-declared by whoever writes the rule, and could in principle be declared inconsistently with what its condition function actually checks. Automatic extraction of requires from unrestricted host-language predicates is out of scope (and undecidable in general, by Rice's theorem); a constrained rule-condition DSL would make it tractable, at the cost of expressiveness.
  • • No human-subject or performance evaluation. All claims about readability improvements over positional Prolog predicates are argued, not measured.
  • • No claim of general-purpose applicability. Karaka roles say nothing about control flow, mutation, or concurrency; this is a declarative/relational calling convention, not a general-purpose language design.

7. Future Work

The limitations above divide cleanly into two kinds: questions whose answers would change what this paper claims, and tasks that need an implementer rather than a result. We separate them accordingly, since only the first kind is research.
Open research questions, in the order we consider most consequential: (1) Does the readability benefit exist and how large is it? A matched-arity comparison of KCC-normalized facts against positional Prolog predicates, measuring comprehension time and error rate across participants, is the single piece of evidence this paper most lacks; the paired materials can be prepared in advance in the repository. (2) What is the right verb-conditioned role mapping? The default vibhakti-Karaka correspondence breaks under passivization and idiosyncratic case government; a per-verb mapping table is structurally identical to a PropBank frameset, which raises the genuinely open question of how much predicate-independence survives contact with real verb semantics. (3) What are query-mode semantics over nested frames? Variables inside sub-events, and quantification across nesting levels, are unspecified beyond one level.
Engineering completion tasks, more mechanical than open: executing generated Cypher against a live graph store rather than emitting text; sandhi-aware input via vidyut-cheda, with the honest complication that segmentation is ambiguous and the parser must surface candidate sets; and dual and plural number.

8. Conclusions

This paper began with a narrow question, whether a fixed vocabulary of grammatical roles could serve as a calling convention, and arrived at a narrower answer than it first expected.
The contribution is a composition, stated after a prior-art search falsified an earlier, broader claim: morphologically-marked Karaka role binding, direct compilation of role-labeled statements to both logic-programming facts and property-graph writes, subsumption-derived rule ordering with compile-time ambiguity rejection, and an assertion/query duality decided by groundness, restricted to declarative/relational programming. The individual components have precedents, now cited; the six-role vocabulary itself appears independently in a contemporaneous system, which we read as evidence of timeliness. The proof-of-concept demonstrates order invariance holds by construction and that the representation compiles to both targets without a lowering pass. Whether the composition yields a measurable readability or maintainability improvement over keyword-argument and named-struct conventions remains the open question this paper explicitly does not settle.

References

  1. Briggs, R. Knowledge representation in Sanskrit and artificial intelligence. AI Mag. 1985, vol. 6(no. 1), 32–39. [Google Scholar]
  2. Fillmore, C. J. The case for case. In Universals in Linguistic Theory; Holt, Rinehart and Winston: New York, 1968. [Google Scholar]
  3. Fillmore, C. J. Frame semantics and the nature of language. Ann. N. Y. Acad. Sci. 1976, vol. 280(no. 1), 20–32. [Google Scholar] [CrossRef]
  4. Minsky, M. A framework for representing knowledge. MIT-AI Laboratory Memo 306, 1974. [Google Scholar]
  5. Banarescu, L.; et al. Abstract meaning representation for sembanking. Proc. 7th Linguistic Annotation Workshop and Interoperability with Discourse, 2013. [Google Scholar]
  6. Kingsbury, P.; Palmer, M. From TreeBank to PropBank. Proc. LREC 2002. [Google Scholar] [CrossRef]
  7. Ernst, M. D.; Kaplan, C.; Chambers, C. Predicate dispatching: a unified theory of dispatch. Proc. ECOOP 1998. [Google Scholar] [CrossRef]
  8. Grosof, B. N. Prioritized conflict handling for logic programs. Proc. Int. Logic Programming Symp., 1997. [Google Scholar]
  9. García, A. J.; Simari, G. R. Defeasible logic programming: an argumentative approach. Theory Pract. Log. Program. 2004, vol. 4(no. 1-2), 95–138. [Google Scholar] [CrossRef]
  10. Khanganba, K. K.; Jha, G. N. Formal Sanskrit syntax: a specification for programming language. Proc. AACL-IJCNLP Student Research Workshop, 2020; pp. 72–78. [Google Scholar]
  11. Prakash, S. Karaka: a Devanagari-first programming language," karaka-core 0.1.0, Rust crates registry. Jun 2026. Available online: https://crates.io/crates/karaka-core.
  12. Huet, G. The Sanskrit Heritage Engine. Available online: https://sanskrit.inria.fr.
  13. Ambuda Project. vidyut: segmentation, sandhi, and Astadhyayi-based word derivation. Available online: https://github.com/ambuda-org/vidyut.
  14. Madathil, K.; et al. sanskrit_parser. Available online: https://github.com/kmadathil/sanskrit_parser.
  15. Implementing Panini's grammar," Language Log, 2023. Available online: https://languagelog.ldc.upenn.edu/nll/?p=61507.
Figure 1. Overview of the Karaka Calling Convention (KCC) Pipeline. A KCC Statement is tokenized and resolved into Karaka roles, assembled into an order independent frame, transformed by ordered sutra rules and compiled to backends.
Figure 1. Overview of the Karaka Calling Convention (KCC) Pipeline. A KCC Statement is tokenized and resolved into Karaka roles, assembled into an order independent frame, transformed by ordered sutra rules and compiled to backends.
Preprints 229772 g001
Figure 2. A frame as a semantic graph. The verb (event) node has outgoing edges labelled with karaka roles. Each role points to its entity (argument). The structure is order independent. .
Figure 2. A frame as a semantic graph. The verb (event) node has outgoing edges labelled with karaka roles. Each role points to its entity (argument). The structure is order independent. .
Preprints 229772 g002
Figure 3. Compilation Targets. A single KCC frame compiles directly to normalized relational facts for Prolog, or property graph writes (like Cypher), since a role labelled relation is already a small labelled edge graph.
Figure 3. Compilation Targets. A single KCC frame compiles directly to normalized relational facts for Prolog, or property graph writes (like Cypher), since a role labelled relation is already a small labelled edge graph.
Preprints 229772 g003
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.