# CodeGraph, A Knowledge Graph of code base

* * *

We've all been there.

You're presenting a technical design. The PowerPoint looks polished. The architecture diagram is clean. The High Level Design (HLD) explains the components, and the Low Level Design (LLD) walks through the implementation.

Then someone asks the inevitable question:

> *"Can you show me where this actually happens in the code?"*

You minimize the presentation, open your IDE, search for a class, jump through imports, open three more files, scroll, explain, backtrack, search again...

Within seconds, the audience has lost the mental model you carefully built on the previous slide.

The problem isn't that our diagrams are wrong.

It's that they're static.

Real software isn't.

A repository is a living graph of modules, imports, services, utilities, and dependencies. Every architectural decision eventually exists as code, but during presentations or even while onboarding a new teammate we're constantly switching between two completely different worlds: beautifully simplified diagrams and thousands of lines of source code.

Neither tells the complete story.

The architecture diagram gives you the *what*.

The code gives you the *how*.

What we really need is something in between.

Imagine clicking on a component in your architecture and immediately seeing the actual code flow behind it. Imagine exploring a repository visually instead of mentally stitching together imports across dozens of files. Better yet, imagine your AI assistant already understanding that structure before you ask your first question.

That idea led to building an AI Skill for Claude with two very different but complementary goals.

* * *

## Two Problems, One Solution

The project intentionally produces two separate artifacts.

**First, something for humans.**

An interactive visualization of the codebase that lets anyone explore the repository in minutes. Instead of navigating folders and jumping between files, you can follow relationships visually. Modules, imports, dependencies, and call paths become part of an explorable map.

It's the missing bridge between architecture diagrams and the source code.

**Second, something for Claude.**

A compact, machine readable structural summary that Claude can query instead of rebuilding its understanding from scratch every session.

If you've used Claude Code on a large repository, you've probably noticed the pattern. Every fresh conversation begins with the same expensive exercise: Claude re opens dozens or sometimes hundreds of files just to reconstruct the same architectural context it already understood yesterday.

![](https://cdn.hashnode.com/uploads/covers/6a50c388383f987664229209/53445413-de4e-4266-a042-ef37fa20c407.png align="center")

Multiply that across every developer and every new session, and an enormous amount of time and tokens are spent rediscovering information that never really changed.

Instead of asking Claude to repeatedly reverse engineer the repository, why not give it a compact structural memory it can consult instantly?

* * *

## One Design Decision Changed Everything

While exploring this idea, one constraint shaped the entire implementation:

**No LLM calls during the indexing pipeline.**

The graph is generated entirely through deterministic static analysis.

Imports are resolved through parsing.

Relationships are extracted using regex based analysis.

The same repository always produces the same graph.

There are no AI generated summaries, no architectural guesses, and no hidden reasoning happening behind the scenes.

At first, this feels like a limitation.

There's no automatically generated description saying *"this service manages user authentication"* or *"this package handles payment processing."*

But the trade off turns out to be surprisingly valuable.

Because there are no LLM calls, generating the graph is almost free.

You can rebuild it after every commit.

You don't worry about token costs.

The graph never drifts because of a hallucinated summary.

And perhaps most importantly, it's completely transparent. Every node and every relationship comes directly from the source code.

* * *

## Why Not Just Use Existing Tools?

This isn't the first attempt to solve the problem.

In my observation there's already an excellent community project called **Understand Anything**, a Claude Code plugin that builds repository knowledge using LLM generated summaries. It's actively maintained and worth exploring. And off course tons of others unknown to me.

But it makes a different set of trade offs.

It spends tokens to understand your repository.

It depends on an LLM to generate that understanding.

And it executes as an external community dependency.

This project deliberately sits at the opposite end of the spectrum.

Instead of richer natural language summaries, it focuses on deterministic structure.

Instead of paying to regenerate knowledge, it can be rebuilt whenever the repository changes.

Instead of asking AI to infer architecture, it extracts what already exists in the code.

Neither approach is universally better, they simply optimize for different priorities.

* * *

## The Bigger Picture

What started as a presentation problem turned into something much more interesting.

The original goal wasn't to build another visualization tool.

It was to eliminate the awkward moment where a technical discussion jumps from a clean architecture slide into a chaotic IDE search.

The visualization helps humans understand large repositories.

The structural summary helps Claude understand them too.

One becomes a living architecture diagram.

The other becomes persistent memory.

Together, they create something we've been missing for years: a shared understanding of a codebase that both developers and AI assistants can navigate without rebuilding context every single time.

In the next sections, we'll dive into how the static analysis works, how the interactive graph is generated, and how the resulting knowledge graph is packaged into an AI Skill that Claude can use as its architectural memory.

* * *

## The Skill Creation

The pipeline is five stages, all static analysis, no model in the loop:

![](https://cdn.hashnode.com/uploads/covers/6a50c388383f987664229209/f2209a0c-4477-4da5-ba7e-3f10c22e0f1b.png align="center")

**Scan** walks the tree, skipping the usual noise `node_modules`, `.git`, build output, vendored dependencies. **Strip** removes comments and string-literal contents before anything else touches the text, which turns out to matter more than it sounds like it should (more on that in a moment). **Extract** pulls imports and top-level function/class/struct definitions with per-language regex patterns. **Resolve** is the hard part: turning a raw import string into an actual file in the repo, whether that means a relative path, a `src/`\-layout absolute import, or a C header reachable only through an include path. **Score** counts inbound edges per file, how many other files depend on it, as a cheap proxy for architectural importance.

Out the other end: a JSON graph for Claude to query, a short Markdown summary for cheap orientation at the start of a session, and a single self-contained HTML dashboard with a force-directed layout, searchable, click-to-inspect that opens in a browser with no server and is safe to screen-share or drop in Slack.

### Making it multi-language: the Redis and nlohmann/json story

Python and JavaScript support came first and were relatively straightforward. C and C++ were not and building them exposed two real bugs that are worth walking through, because they're the kind of thing that only shows up once you stop testing against toy examples and point the tool at real code.

Testing against **Redis** (217 C files) surfaced the first one. A comment in `util.c` read: *"Based on the following article (that apparently does not provide a novel approach...)"* and the regex meant to detect function signatures matched `article` as a function name, then greedily consumed everything up to the *next* real closing parenthesis and opening brace it could find, which belonged to the actual function that followed. The comment didn't just create a fake entry it ate the real one. `ull2string`, a genuine function, silently disappeared from the graph.

The fix was to strip comments before running any extraction. The first attempt at that fix introduced a second bug: it also blanked the *contents* of string literals, which included the quoted header paths inside `#include "server.h"`. Edge count on the same Redis run went from 525 to 3. The second fix separated the two concerns strip comments before both imports and definitions, but only blank string contents before definition-extraction, since import-extraction needs those quotes intact.

Testing against **nlohmann/json** (a template-heavy C++ header library) found a third: template parameter lists like `template<class ArrayType = std::vector<int>>` were being read as class definitions named `ArrayType`. A negative lookahead fixed it and immediately introduced a fourth bug, truncating real class names (`Allocator` → `Allocato`) by letting the regex backtrack mid-identifier. A word-boundary anchor closed that gap.

What both real repos validated in the end:

| Repo | Language | Files | Edges | What it got right |
| --- | --- | --- | --- | --- |
| Redis | C | 217 | 525 | Correctly flagged `server.h` as the most depended-on file which is genuinely true of Redis's architecture |
| nlohmann/json | C++ | 47 | 160 | Correctly extracted `basic_json`, `lexer`, `parser` as real classes, zero template-parameter noise |
| Flask | Python | 83 | 211 | Regression-checked after every C/C++ change to confirm nothing broke |

![](https://cdn.hashnode.com/uploads/covers/6a50c388383f987664229209/dfd419d0-9c9e-434b-b507-ad4d266bf867.png align="center")
