--- license: other license_name: mii-open-license-v1.1 license_link: https://github.com/cripto-bot/graphlang/blob/main/LICENSE language: - en tags: - intermediate-representation - semantic-ir - code-analysis - compiler - cross-language - ast - code-compression - dataset - demo pipeline_tag: other --- # GraphLang — Universal Semantic Kernel for Code **The same computational intent, in 13 languages, collapses to the same 12-node graph.** GraphLang is a semantic Intermediate Representation (IR) that maps source code from 13 languages (Python, Java, JavaScript, TypeScript, C#, Rust, Go, Kotlin, Ruby, PHP, Zig, C, C++) into a single canonical graph of **12 universal IR kinds**. It is not a new language — it is a discovery: different syntaxes converge to the same structure when their intent is equivalent. > **[Try it now on Spaces →](https://huggingface.co/spaces/Jose-dev/graphlang)** > — paste code and see it collapse into a GraphLang IR graph, live. ```text Python: def add(a, b): return a + b ─┐ Java: int add(int a, int b){ return a+b; } ─┤ → SAME GraphLang IR JavaScript: function add(a,b){ return a+b; } ─┘ (identical graph) ``` ## The 12 IR Kinds (FROZEN) | # | Kind | Meaning | |---|------|---------| | 1 | `function` | Executable unit with parameters | | 2 | `if` | Conditional branch | | 3 | `for` | Bounded iteration | | 4 | `while` | Unbounded iteration | | 5 | `return` | Value return | | 6 | `assign` | Variable binding | | 7 | `call` | Invocation | | 8 | `binop` | Binary / comparison operation | | 9 | `unary` | Unary operation | | 10 | `var` | Variable reference | | 11 | `const` | Literal constant | | 12 | `block` | Statement sequence | The 12 kinds were derived from the analysis of ~2,215 Concrete Syntax Tree (CST) node types across the 13 languages. ## Cross-language equivalence (measured, not claimed) 12 canonical intents were compiled in Python, Java, and JavaScript through the **real engine** (`complete_normalizer.py`) and fingerprinted with a recursive structural hash (content-addressed, not node-id based). | Group | Identical IR? | Notes | |-------|---------------|-------| | `add`, `max2`, `abs`, `square`, `compare`, `negate`, `twice` | **Yes (7/7)** | arithmetic / expression intents | | `is_even`, `sum_loop`, `fact`, `fib`, `grade` | Partial | comparison/loop sugar (`elif` vs `else if`, `for-of` vs `for-in`, Python `Compare` vs binary op) | Arithmetic and expression-level intents collapse to **bit-identical IR** across the three languages. The remaining divergence is a known normalization gap in comparison and loop sugar — documented, not hidden. See `data/equivalence.csv`. ## IR visualization `grade(score)` in Python, normalized to GraphLang IR: ![GraphLang IR graph](assets/grade.png) ## Reproducible Benchmarks The compression ratio converges to a constant — **22.5x monolingual** and **29.8x multilingual** — from 100K functions onward. Results below are reproducible with the engine (`benchmark_100k.py` / `benchmark_1m.py`). | Functions | Total Nodes | Unique Patterns | Ratio | Errors | |-----------|-------------|-----------------|-------|--------| | 1,500 | 32,481 | 1,567 | 20.7x | 0 | | 10,000 | 217,233 | 9,770 | 22.2x | 0 | | 100,000 | 2,170,018 | 96,616 | 22.5x | 0 | | 1,000,000 | 21,721,197 | 965,045 | 22.5x | 0 | ## Datasets (downloadable) | File | Content | |------|---------| | `data/examples.jsonl` | 36 code→IR samples (12 intents × Python/Java/JavaScript), full IR node dump + fingerprint | | `data/equivalence.csv` | Cross-language equivalence matrix with recursive structural hashes | | `benchmark_100k_results.json` | 100K-function compression measurement | | `benchmark_1m_results.json` | 1M-function compression measurement | ## Included in this repository - `app.py` — **didactic** interactive demo (Gradio): paste code, see the IR graph, merge two functions, measure structural deduplication. This is a self-contained educational reimplementation, **not** the production engine. - `parallel_ir.py` — GPU/HPC extension (CUDA / OpenCL / Metal detection and thread-index normalization) sitting on top of the 12 core kinds. - `benchmark_100k_results.json`, `benchmark_1m_results.json` — reproducible benchmark measurements. - `data/` — the code→IR dataset and equivalence matrix (generated by the real engine). - `assets/` — IR graph visualizations (SVG + PNG). - `paper/` — the academic paper (GraphLang: a universal semantic kernel for code). ## Engine and license The full multi-language normalizer engine is available under the **MII Open License v1.1** (see `LICENSE`): free for non-commercial and research use, restricted for AI/ML training (drafted as a copyright condition, enforceable under *Jacobsen v. Katzer*), and commercial use requires a license. The benchmarks and `data/` files in this repository were produced by the real engine. The `app.py` demo is a simplified didactic reimplementation that illustrates the concept only. For the engine, the 20M aligned function-pair dataset, or commercial licensing: **josu31.jas@gmail.com** - Source repository: - Author: **Josué Argaña Silguero** — 2026 --- *"No inventamos un nuevo lenguaje. Descubrimos que todos los lenguajes ya hablaban el mismo."*