cognee-ingestion
Use when putting data into cognee memory with remember() — choosing inputs (text, files, folders, URLs, repos, databases), datasets and node_sets, loaders, ontologies, the graph extractor (LLM or GLiNER), chunking, dry-run cost estimates, or when remember() raises on a keyword argument.
Install / Use
npx skills add topoteretes/cognee --skill cognee-ingestionInstalls into whichever agent you are using.
SKILL.md
Installable skill definition
Quality Score
Category
AI & Machine LearningSupported Platforms
Our assessment of cognee-ingestion
cognee-ingestion scores 96/100 on our quality scale, 80th of 951 AI & Machine Learning skills we index (top 9%).
Its SKILL.md is 10 KB long, well organised into 13 sections with 3 code examples: a thorough specification that gives an agent plenty to work with.
With 30,958 GitHub stars, it is one of the more widely adopted skills in the catalogue.
Maintenance, license and trust
- The repository was last updated 13 days ago, so cognee-ingestion is actively maintained.
- It is released under the Apache-2.0 license, a permissive license that allows use, modification and commercial use with attribution.
- Its trust signals score 100/100, with no cautions. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.
cognee-ingestion compared with similar skills
All 4 of these similar skills score higher than cognee-ingestion; compare them before choosing.
| Skill | Score | Stars | Updated | Format |
|---|---|---|---|---|
| cognee-ingestion (this skill)by topoteretes | 96 | 31.0k | 13d ago | SKILL.md |
| claude-memby thedotmack | 100 | 97.7k | 1d ago | CLAUDE.md |
| Agent-Reachby Panniantong | 100 | 93.2k | today | CLAUDE.md |
| Understand-Anythingby Egonex-AI | 100 | 85.5k | 2d ago | CLAUDE.md |
| headroomby headroomlabs-ai | 100 | 74.6k | today | CLAUDE.md |
Frequently asked questions
- How do I install cognee-ingestion?
- Run
npx skills add topoteretes/cognee --skill cognee-ingestion. The install tabs above show the steps for each supported agent. - Which AI agents does cognee-ingestion work with?
- It is written for Universal, as a SKILL.md file. Other agents that read the same format can often use it too.
- Is cognee-ingestion safe to use?
- It is Apache-2.0-licensed and scores 100/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
- Is cognee-ingestion still maintained?
- The repository was last updated 13 days ago, so cognee-ingestion is actively maintained.
Skill content
View source on GitHubname: cognee-ingestion description: Use when putting data into cognee memory with remember() — choosing inputs (text, files, folders, URLs, repos, databases), datasets and node_sets, loaders, ontologies, the graph extractor (LLM or GLiNER), chunking, dry-run cost estimates, or when remember() raises on a keyword argument.
Ingest data with remember()
remember() is cognee's ingestion API. One call stores the data, builds the
knowledge graph, and enriches it. Use it for all ingestion; every option in
this skill is a remember() argument unless it says otherwise.
import cognee
result = await cognee.remember("Einstein was born in Ulm.") # text
result = await cognee.remember(
["./notes.md", "./report.pdf"], # files
dataset_name="research",
)
print(result.status, result.dataset_id) # "completed", UUID
All cognee functions are async. Without dataset_name data goes to
main_dataset. Needs LLM_API_KEY unless you use the GLiNER extractor
(below).
Use it
Inputs
data accepts a string, a list of strings, file paths (absolute, file://,
s3://), http(s) URLs, binary streams, or a list mixing them.
- URLs are fetched and scraped (needs
ALLOW_HTTP_REQUESTS=true, the default). - Folders are ingested file by file. A folder that looks like a code
project, or a GitHub/GitLab URL, becomes one code repository (needs
giton PATH). - Code files (
.py,.ts,.go, …) go down the code-graph route: a deterministic graph, no LLM calls, searchable only withSearchType.CODE. To index a whole repository explicitly, passcontent_type="code". - Databases and dlt sources: a SQL connection string, a dlt
DltResource/DltSource, or a CSV. dlt is a core dependency, so no extra is needed (cognee[dlt]is an empty compatibility extra). Options:primary_key(default"id"),write_disposition("replace"default, or"append"),query,max_rows_per_table. - Skill playbooks (
SKILL.mdfiles):content_type="skills"; ingests into the target dataset (defaultmain_dataset), so passdataset_nameto keep skills in their own dataset.
Where the data goes
| Argument | What it does |
|---|---|
| dataset_name / dataset_id | Target dataset. dataset_id wins. A dataset is the unit of permissions and isolation. |
| node_set=["AI", "FinTech"] | Tags the data so recall can filter to it later with recall(..., node_name=["AI"]). |
| session_id="chat_1" | Writes to the fast session cache instead of the graph; improve() bridges it into the graph in the background. See the cognee-improve-sessions skill. Requires CACHING=true. |
How the graph is built
| Argument | What it does |
|---|---|
| extractor | "llm" or "gliner_demo" (alias "gliner"). Default is GRAPH_EXTRACTOR=auto: the LLM when an API key is configured, otherwise GLiNER. |
| graph_model=MyModel | Extract into your own DataPoint model instead of the generic KnowledgeGraph. See the cognee-custom-graph-models skill. |
| custom_prompt | Replaces the entity-extraction prompt (ignored by GLiNER). |
| config={"ontology_config": {...}} | Ground entities in an OWL ontology (below). |
| chunk_size, chunker | Max tokens per chunk (default: derived from the embedding and LLM limits) and the chunker class (default TextChunker). |
| preferred_loaders | Choose a loader per file type (below). |
| self_improvement | Default True: runs improve() after the graph is built. Its outcome is on result.improve / result.improve_error; a failed improve never fails the remember. |
| run_in_background=True | Returns immediately with status="running"; await result to wait. |
Ontologies
from cognee.modules.ontology.rdf_xml.RDFLibOntologyResolver import RDFLibOntologyResolver
config = {
"ontology_config": {
"ontology_resolver": RDFLibOntologyResolver(ontology_file="./my.owl"),
# "ontology_mode": "strict", # drop entities with no ontology match
}
}
await cognee.remember(texts, config=config)
Or set ONTOLOGY_FILE_PATH (plus ONTOLOGY_MODE, MATCHING_STRATEGY) in
.env. annotate (default) only enriches; strict drops entities that
match no ontology class or individual. It prunes only the graph, chunk text
is still stored. Strict mode with an empty or missing ontology file is a hard
error. Over HTTP, upload the ontology to /api/v1/ontologies and pass its
ontology_key to POST /api/v1/remember. Example:
examples/guides/ontology_quickstart.py.
Loaders
Each file is claimed by the first loader that accepts it. Default order:
code, text, pypdf, image, audio, video, dlt_csv, csv, unstructured,
advanced_pdf, docling. Names: text_loader, code_loader, csv_loader,
dlt_csv_loader, pypdf_loader, image_loader, audio_loader,
video_loader, unstructured_loader, advanced_pdf_loader,
docling_loader, beautiful_soup_loader.
# Treat a code file as a plain document (chunking + LLM extraction):
await cognee.remember("./script.py", preferred_loaders={"text_loader": {}})
Office formats (DOCX, PPTX, …) need the docs (unstructured) or docling
extra. A preferred loader that is not installed is skipped with only an info
log, so check the extra is installed when a file comes out wrong.
Check the cost first
dry_run=True returns a token and cost estimate without ingesting anything
or calling the LLM. It excludes the calls improve() makes. Not supported
with GLiNER, sessions, or a remote instance.
dry_run="presort" on a folder returns a PresortReport (junk, duplicates,
version candidates, possible personal data, proposed dataset groups). Apply
it with await cognee.remember(report), or pass auto_apply=True.
Without an LLM: GLiNER
extractor="gliner" builds the graph and summaries with a local GLiNER2
model, with no LLM call (embeddings still run). Install
pip install "cognee[gliner]"; the model (about 750 MB) downloads on first use.
It cannot be combined with a custom graph_model, dry_run,
session_id, or a remote instance.
For production: the open-source GLiNER extractor is a demo. cognee's enterprise GLiNER extraction is more accurate and covers more labels. The same goes for the Postgres graph adapter (
postgres_demo). Contact social@cognee.ai.
Pitfalls
-
Unknown keyword arguments raise.
remember()forwards kwargs through a fixed allow-list and raisesTypeError: Unexpected keyword argumentsfor anything else. These real options are not on it yet:| Option | Workaround through remember() | |---|---| |
ontology_file_path|config={"ontology_config": ...}orONTOLOGY_FILE_PATH(above) | |functional_relationships,chunk_attachment| None yet. Onlycognee.cognify()accepts them. | |extraction_rules| Pass it through the loader:preferred_loaders={"beautiful_soup_loader": {"extraction_rules": {...}}}(works inremember()andadd()). Needs thescrapingextra: without it the loader is not registered and the rules are silently ignored | |tavily_config,soup_crawler_config| Not honoured byadd()orremember(); only thecognee/tasks/web_scrapertasks use them | |column_value_columns(dlt) | None yet. Onlycognee.add()accepts it. |If a user needs one with no workaround, say so plainly: the option exists on the lower-level
add()/cognify()but not onremember()yet. -
Changed files raise
DocumentUpdateRequiredError. Re-remembering the same path (or the same filename for an upload) with different content is an update, not a new document. Usecognee.update(data_id=..., data=..., dataset_id=...), which re-extracts only the changed chunks and keeps the document's id. Identical content is a no-op. -
content_typeis strict. OnlyNone,"skills", or"code"."code"rejectssession_idand needs repository paths or git URLs;"skills"ingests into the target dataset like any other call (defaultmain_dataset); passdataset_nameto keep skills in their own dataset. -
Session mode needs
CACHING=true, andextractorcannot be combined withsession_id. -
Remote mode. After
cognee.serve(url), calls go to the server:extractorandsession_idsraise, and other options the client does not forward (includinggraph_model,node_set, and ontologyconfig) are dropped without an error. -
Every remember runs
improve()unlessself_improvement=FalseorIMPROVE_AUTO_ENABLED=false. In scripts, callawait cognee.wait_for_background_tasks()before exiting.
How it works
remember(data) runs add() (store raw data and create Data rows), then
cognify() (classify documents, chunk, extract the graph and summaries,
store in graph and vector DBs), then improve(). remember(data, session_id=...) writes to the session cache instead.
- Entry point and kwarg routing:
cognee/api/v1/remember/remember.py(RememberKwargs,_ADD_ONLY/_COGNIFY_ONLY/_SHARED) - Storage:
cognee/api/v1/add/add.py,cognee/tasks/ingestion/ingest_data.py - Graph build:
cognee/api/v1/cognify/cognify.py,cognee/tasks/graph/extract_graph_from_data.py,cognee/tasks/storage/add_data_points.py - Extractor choice:
cognee/modules/cognify/config.py:resolve_extractor; GLiNER package:cognee/tasks/graph/gliner_demo/ - Ontologies:
cognee/modules/ontology/ - Loaders:
cognee/infrastructure/loaders/(supported_loaders.py,LoaderEngine.py) - dlt:
cognee/tasks/ingestion/resolve_dlt_sources.py
Examples in examples/guides/: simple_cognee_example.py,
nodeset_grouping_example.py, ontology_quickstart.py,
gliner_demo_llm_free_cognify.py, no_llm_remember_recall.py,
temporal_recall.py, presort_downloads.py,
web_url_content_ingestion_example.py, code_graph_example.py.
Extending it
- New remember() option: add it to
RememberKwargsand to the matching routing set inremember.py. An option onadd()/cognify()that is not in a routing set raisesTypeErrorfromremember(). - New loader: implement
LoaderInterface(cognee/infrastructure/loaders/LoaderInterface.py), register it insupported_loaders.py(extras-gated loaders go underexternal/), and add it to the priority list inLoaderEngine.pyif it should run by default. - New cognify task: see the
cognee-custom-pipelinesskill andcognee/tasks/README.md.
Related Skills
claude-mem
97.7kPersistent Context Across Sessions for Every Agent – Captures everything your agent does during sessions, compresses it with AI, and injects relevant context back into future sessions. Works with Claude Code, OpenClaw, Codex, Gemini, Hermes, Copilot, OpenCode + More
Agent-Reach
93.2kGive your AI agent eyes to see the entire internet. Read & search Twitter, Reddit, YouTube, GitHub, Bilibili, XiaoHongShu — one CLI, zero API fees.
Understand-Anything
85.5kGraphs that teach > graphs that impress. Turn any code into an interactive knowledge graph you can explore, search, and ask questions about. Works with Claude Code, Codex, Cursor, Copilot, Gemini CLI, and more.
headroom
74.6kCompress tool outputs, logs, files, and RAG chunks before they reach the LLM. 20% fewer tokens for coding agents, 60-95% fewer tokens for JSON, same answers. Library, proxy, MCP server.
Languages
Trust signals
From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.
