cognee-custom-graph-models
Use when defining the shape of cognee's knowledge graph with graph_model= — writing DataPoint node classes, choosing identity and index fields so nodes merge and are searchable, declaring typed Edge fields and FromIdentity references, building a model from a JSON schema, or debugging duplicated node…
Install / Use
npx skills add topoteretes/cognee --skill cognee-custom-graph-modelsInstalls into whichever agent you are using.
SKILL.md
Installable skill definition
Quality Score
Category
AI & Machine LearningSupported Platforms
Tags
Our assessment of cognee-custom-graph-models
cognee-custom-graph-models scores 94/100 on our quality scale, 142nd of 951 AI & Machine Learning skills we index (top 15%).
Its SKILL.md is 9.6 KB long, well organised into 11 sections with 1 code example: a thorough specification that gives an agent plenty to work with.
With 30,958 GitHub stars, it is one of the more widely adopted skills in the catalogue.
Maintenance, license and trust
- The repository was last updated 13 days ago, so cognee-custom-graph-models is actively maintained.
- It is released under the Apache-2.0 license, a permissive license that allows use, modification and commercial use with attribution.
- Its trust signals score 100/100, with no cautions. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.
cognee-custom-graph-models compared with similar skills
All 4 of these similar skills score higher than cognee-custom-graph-models; compare them before choosing.
| Skill | Score | Stars | Updated | Format |
|---|---|---|---|---|
| cognee-custom-graph-models (this skill)by topoteretes | 94 | 31.0k | 13d ago | SKILL.md |
| claude-memby thedotmack | 100 | 97.7k | 1d ago | CLAUDE.md |
| Understand-Anythingby Egonex-AI | 100 | 85.5k | 2d ago | CLAUDE.md |
| headroomby headroomlabs-ai | 100 | 74.6k | today | CLAUDE.md |
| CowAgentby zhayujie | 100 | 47.3k | today | CLAUDE.md |
Frequently asked questions
- How do I install cognee-custom-graph-models?
- Run
npx skills add topoteretes/cognee --skill cognee-custom-graph-models. The install tabs above show the steps for each supported agent. - Which AI agents does cognee-custom-graph-models work with?
- It is written for Universal, as a SKILL.md file. Other agents that read the same format can often use it too.
- Is cognee-custom-graph-models safe to use?
- It is Apache-2.0-licensed and scores 100/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
- Is cognee-custom-graph-models still maintained?
- The repository was last updated 13 days ago, so cognee-custom-graph-models is actively maintained.
Skill content
View source on GitHubname: cognee-custom-graph-models description: Use when defining the shape of cognee's knowledge graph with graph_model= — writing DataPoint node classes, choosing identity and index fields so nodes merge and are searchable, declaring typed Edge fields and FromIdentity references, building a model from a JSON schema, or debugging duplicated nodes, missing edges, or InvalidReferenceTypeError.
Custom graph models
By default cognee extracts a generic KnowledgeGraph of entities and
relationships. Pass your own model with graph_model= and the LLM fills
your node and edge types instead.
from typing import Annotated, Literal
import cognee
from cognee.low_level import DataPoint, Edge, FromIdentity
class Role(DataPoint):
name: str
metadata: dict = {"index_fields": ["name"], "identity_fields": ["name"]}
class Person(DataPoint):
name: str
is_a: Annotated[Role, FromIdentity()] | None = None # reference by name
reports_to: list[Edge["Person", "Person"]] = [] # edge owned by Person
metadata: dict = {"index_fields": ["name"], "identity_fields": ["name"]}
class PeopleGraph(DataPoint): # the root the LLM fills
people: list[Person]
friends_with: list[Edge[Person, Person]] = []
family: list[Edge[Person, Person, Literal["married_to", "sibling_of"]]] = []
metadata: dict = {"index_fields": [], "transparent": True}
await cognee.remember(text, graph_model=PeopleGraph, custom_prompt="Extract every person...")
Full example: examples/guides/custom_graph_model.py.
Use it
Nodes: DataPoint classes
Every node type subclasses DataPoint (from cognee.low_level import DataPoint). Its fields become:
- Node properties: scalars, strings, dicts, and anything that is not a DataPoint.
- Edges named after the field: a field holding a DataPoint or a list of
DataPoints.
members: list[Person]becomesmembersedges.
dict[str, Person], sets and plain tuples are stored as properties, not
edges.
Identity and search: metadata
| Key | What it does |
|---|---|
| identity_fields | The node id is derived from these field values (normalized: lowercased, spaces to _, apostrophes removed). The same entity from two chunks or two runs becomes one node. |
| index_fields | Each field gets a vector collection named <ClassName>_<field>, so recall can find the node. |
| transparent | The node is not stored; its children take its place. Use it for a root container like PeopleGraph. |
Without identity_fields every node gets a random id, so the same person
is duplicated in every chunk and every run. Set it on every node type that
represents a real-world entity.
Write metadata explicitly, as in the examples above. There is also an
annotation shortcut (from cognee.infrastructure.engine import Dedup, Embeddable; name: Annotated[str, Embeddable(), Dedup()]), but today only
half of it works:
Dedup()works: ids are derived from the marked fields.Embeddable()does not index. The markers update the class-level default, but each instance still carries{"index_fields": []}, and indexing reads the instance, so no vector collection is created and recall cannot find the node.
Markers are also ignored entirely when the class declares metadata
itself.
Typed edges: list[Edge[Source, Target, Name]]
The LLM answers edges as flat rows of identity strings (source, target),
and cognee resolves them to the extracted nodes. The third parameter
controls the relationship name:
| Declaration | Relationship name |
|---|---|
| list[Edge[Person, Person]] | The field name (friends_with) |
| list[Edge[Person, Person, Literal["a", "b"]]] | The LLM picks one value |
| list[Edge[Person, Person, str]] | Free-form from the LLM, normalized |
- Where to declare: on the root model for relationships with no obvious
owner, or on the owning node. On the owner, endpoints of the owner's own
type must be strings (
Edge["Person", "Person"]), because the class is not defined yet inside its own body. - Always a list:
Edge[...]orEdge[...] | Noneon its own raises. - Both endpoint types need exactly one
identity_fieldsentry.
References: Annotated[Target, FromIdentity()]
Instead of a nested object, the LLM answers the identity string of a node
(is_a: "engineer"), and cognee links to that node. Supported spellings:
Target, Target | None, list[Target], list[Target] | None.
Anything else raises InvalidReferenceTypeError. The target needs exactly
one identity field, and its other required fields need defaults.
Edge values you build by hand
Edge(source=..., target=..., relationship_type=..., weight=..., properties={...}). An omitted source falls back to the node declaring the
field; on a parametrized field that node must be the declared Source
type, or it raises. On a root container, always pass source=. The tuple
form (Edge(weight=0.8), target_node) attaches edge properties to a plain
DataPoint field (examples/guides/custom_data_models.py).
From JSON instead of Python
cognee.low_level.graph_model_from_spec(spec): a small entity/relation spec (names, fields,one/manyrelations) compiled to DataPoint classes, with identity and index onnameby default. Example:examples/guides/graph_model_from_json.py.cognee.low_level.graph_schema_to_graph_model(json_schema): a JSON Schema (needs a top-leveltitle; only internal#refs).- HTTP:
POST /api/v1/remembertakes agraph_modelform field (JSON schema string);POST /api/v1/cognifytakes agraphModelJSON object.POST /api/v1/llm/infer-schemaproposes a schema from sample text.
Neither JSON path can express typed Edge fields or FromIdentity; use
Python classes for those.
Pitfalls
- Duplicated nodes → missing
identity_fields. - Node never shows up in recall → no
index_fields, or the recalling process never imported the model class (graph completion searches the collections of DataPoint classes loaded in that process). - Edge rows silently missing → an unresolved row is dropped with a warning. Endpoints resolve by exact type against nodes extracted from the same chunk: a subclass instance does not match where its base is declared, and a node from another chunk or an earlier run is not a candidate.
InvalidReferenceTypeError, "declares Edge in a shape the LLM extraction cannot fill" → an unsupportedFromIdentityorEdgespelling. These are raised when the model is converted during extraction, not at class definition, so they appear mid-pipeline.- String endpoints resolve only to the owning model itself or a
module-level class. A string naming another class defined inside a
function raises
InvalidReferenceTypeError, so define models at module level. - Field names that collide with DataPoint's own fields (
id,type,version,metadata,created_at,belongs_to_set, …) are stripped from what the LLM sees. Rename them. - A subclass that overrides
metadatareplaces the parent's entirely. Droppingidentity_fieldsonly logs a warning. A subclass also hashes ids under its own class name. - Write a
custom_prompt. Without one the generic knowledge-graph prompt is used; your schema reaches the LLM only as structured output. - Not with GLiNER.
extractor="gliner_demo"(alias"gliner") raises with a customgraph_model, and so does the defaultGRAPH_EXTRACTOR=autowhen no LLM key is configured (it resolves togliner_demo). - Remote mode drops it. After
cognee.serve(url),remember()andcognify()do not forwardgraph_model; the server builds a generic graph. - A custom model skips the generic path's ontology resolution, per-graph
node dedup, and
functional_relationships. Summaries still run.
How it works
extract_content_graph converts a DataPoint model into a plain Pydantic
model for the LLM: infrastructure fields and metadata are stripped, typed
edge fields become row lists (FriendsWithEdge with source/target
strings), and FromIdentity fields become strings. The answer is converted
back into DataPoint instances with ids from identity_fields, edge rows are
resolved against the nodes in that answer, and the result is attached to the
chunk (chunk.contains) and stored by add_data_points. Ownership is
recorded per document, so forget(data_id=...) removes a custom-model
document's nodes while shared nodes survive.
- LLM boundary, both directions:
cognee/shared/llm_graph_model.py - DataPoint, metadata, ids:
cognee/infrastructure/engine/models/DataPoint.py - Markers:
cognee/infrastructure/engine/models/FieldAnnotations.py - Edge:
cognee/infrastructure/engine/models/Edge.py - Property vs edge decision:
cognee/modules/graph/utils/field_edges.py - Custom-model branch of extraction:
cognee/tasks/graph/extract_graph_from_data.py - JSON schema / spec paths:
cognee/shared/graph_model_utils.py,cognee/modules/graph_models/ - Vector collections:
cognee/tasks/storage/index_data_points.py
Extending it
- Tests for the LLM round trip:
cognee/tests/unit/modules/graph/test_content_graph_to_data_point.py. Edge typing:cognee/tests/unit/interfaces/graph/test_typed_edge_model.py,test_typed_edges_graph.py. Identity:cognee/tests/unit/infrastructure/engine/test_identity_fields.py. Property vs edge:cognee/tests/unit/modules/graph/test_field_edges.py. - Deletion of custom-model nodes:
cognee/tests/test_delete_custom_graph.py. - A new
EdgeorFromIdentityspelling must be handled in both directions inllm_graph_model.py, and rejected withInvalidReferenceTypeErrorwhen unsupported, never silently accepted.
Related Skills
claude-mem
97.7kPersistent Context Across Sessions for Every Agent – Captures everything your agent does during sessions, compresses it with AI, and injects relevant context back into future sessions. Works with Claude Code, OpenClaw, Codex, Gemini, Hermes, Copilot, OpenCode + More
Understand-Anything
85.5kGraphs that teach > graphs that impress. Turn any code into an interactive knowledge graph you can explore, search, and ask questions about. Works with Claude Code, Codex, Cursor, Copilot, Gemini CLI, and more.
headroom
74.6kCompress tool outputs, logs, files, and RAG chunks before they reach the LLM. 20% fewer tokens for coding agents, 60-95% fewer tokens for JSON, same answers. Library, proxy, MCP server.
CowAgent
47.3kOpen-source personal AI assistant & Agent Harness. Plans tasks, runs tools and skills, self-evolves with memory and knowledge. Multi-agent, multi-model, multi-channel. Lightweight, extensible, one-line install.
Languages
Trust signals
From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.
