docx
Word document manipulation with python-docx - handling split placeholders, headers/footers, nested tables
Install / Use
npx skills add benchflow-ai/skillsbench --skill docxInstalls into whichever agent you are using.
SKILL.md
Installable skill definition
Quality Score
Category
Content & MediaSupported Platforms
Our assessment of docx
docx scores 89/100 on our quality scale, 465th of 1,214 Content & Media skills we index (top 39%).
Its SKILL.md is 7.7 KB long, well organised into 13 sections with 7 code examples: a thorough specification that gives an agent plenty to work with.
With 1,813 GitHub stars, it is one of the more widely adopted skills in the catalogue.
Maintenance, license and trust
- The repository was last updated about 2 months ago, so docx is actively maintained.
- It is released under the Apache-2.0 license, a permissive license that allows use, modification and commercial use with attribution.
- Its trust signals score 100/100, with no cautions. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.
docx compared with similar skills
All 4 of these similar skills score higher than docx; compare them before choosing.
| Skill | Score | Stars | Updated | Format |
|---|---|---|---|---|
| docx (this skill)by benchflow-ai | 89 | 1.8k | 2mo ago | SKILL.md |
| Agent-Reachby Panniantong | 100 | 90.5k | 19d ago | CLAUDE.md |
| headroomby headroomlabs-ai | 100 | 74.4k | today | CLAUDE.md |
| Scraplingby D4Vinci | 100 | 85.6k | today | MCP Server |
| crawl4aiby unclecode | 100 | 84.8k | 9d ago | MCP Server |
Frequently asked questions
- How do I install docx?
- Run
npx skills add benchflow-ai/skillsbench --skill docx. The install tabs above show the steps for each supported agent. - Which AI agents does docx work with?
- It is written for Universal, as a SKILL.md file. Other agents that read the same format can often use it too.
- Is docx safe to use?
- It is Apache-2.0-licensed and scores 100/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
- Is docx still maintained?
- The repository was last updated about 2 months ago, so docx is actively maintained.
Skill content
View source on GitHubname: docx description: Word document manipulation with python-docx - handling split placeholders, headers/footers, nested tables
Word Document Manipulation with python-docx
Critical: Split Placeholder Problem
The #1 issue with Word templates: Word often splits placeholder text across multiple XML runs. For example, {{CANDIDATE_NAME}} might be stored as:
- Run 1:
{{CANDI - Run 2:
DATE_NAME}}
This happens due to spell-check, formatting changes, or Word's internal XML structure.
Naive Approach (FAILS on split placeholders)
# DON'T DO THIS - won't find split placeholders
for para in doc.paragraphs:
for run in para.runs:
if '{{NAME}}' in run.text: # Won't match if split!
run.text = run.text.replace('{{NAME}}', value)
Correct Approach: Paragraph-Level Search and Rebuild
import re
def replace_placeholder_robust(paragraph, placeholder, value):
"""Replace placeholder that may be split across runs."""
full_text = paragraph.text
if placeholder not in full_text:
return False
# Find all runs and their positions
runs = paragraph.runs
if not runs:
return False
# Build mapping of character positions to runs
char_to_run = []
for run in runs:
for char in run.text:
char_to_run.append(run)
# Find placeholder position
start_idx = full_text.find(placeholder)
end_idx = start_idx + len(placeholder)
# Get runs that contain the placeholder
if start_idx >= len(char_to_run):
return False
start_run = char_to_run[start_idx]
# Clear all runs and rebuild with replacement
new_text = full_text.replace(placeholder, str(value))
# Preserve first run's formatting, clear others
for i, run in enumerate(runs):
if i == 0:
run.text = new_text
else:
run.text = ''
return True
Best Practice: Regex-Based Full Replacement
import re
from docx import Document
def replace_all_placeholders(doc, data):
"""Replace all {{KEY}} placeholders with values from data dict."""
def replace_in_paragraph(para):
"""Replace placeholders in a single paragraph."""
text = para.text
# Find all placeholders
pattern = r'\{\{([A-Z_]+)\}\}'
matches = re.findall(pattern, text)
if not matches:
return
# Build new text with replacements
new_text = text
for key in matches:
placeholder = '{{' + key + '}}'
if key in data:
new_text = new_text.replace(placeholder, str(data[key]))
# If text changed, rebuild paragraph
if new_text != text:
# Clear all runs, put new text in first run
runs = para.runs
if runs:
runs[0].text = new_text
for run in runs[1:]:
run.text = ''
# Process all paragraphs
for para in doc.paragraphs:
replace_in_paragraph(para)
# Process tables (including nested)
for table in doc.tables:
for row in table.rows:
for cell in row.cells:
for para in cell.paragraphs:
replace_in_paragraph(para)
# Handle nested tables
for nested_table in cell.tables:
for nested_row in nested_table.rows:
for nested_cell in nested_row.cells:
for para in nested_cell.paragraphs:
replace_in_paragraph(para)
# Process headers and footers
for section in doc.sections:
for para in section.header.paragraphs:
replace_in_paragraph(para)
for para in section.footer.paragraphs:
replace_in_paragraph(para)
Headers and Footers
Headers/footers are separate from main document body:
from docx import Document
doc = Document('template.docx')
# Access headers/footers through sections
for section in doc.sections:
# Header
header = section.header
for para in header.paragraphs:
# Process paragraphs
pass
# Footer
footer = section.footer
for para in footer.paragraphs:
# Process paragraphs
pass
Nested Tables
Tables can contain other tables. Must recurse:
def process_table(table, data):
"""Process table including nested tables."""
for row in table.rows:
for cell in row.cells:
# Process paragraphs in cell
for para in cell.paragraphs:
replace_in_paragraph(para, data)
# Recurse into nested tables
for nested_table in cell.tables:
process_table(nested_table, data)
Conditional Sections
For {{IF_CONDITION}}...{{END_IF_CONDITION}} patterns:
def handle_conditional(doc, condition_key, should_include, data):
"""Remove or keep conditional sections."""
start_marker = '{{IF_' + condition_key + '}}'
end_marker = '{{END_IF_' + condition_key + '}}'
for para in doc.paragraphs:
text = para.text
if start_marker in text and end_marker in text:
if should_include:
# Remove just the markers
new_text = text.replace(start_marker, '').replace(end_marker, '')
# Also replace any placeholders inside
for key, val in data.items():
new_text = new_text.replace('{{' + key + '}}', str(val))
else:
# Remove entire content between markers
new_text = ''
# Apply to first run
if para.runs:
para.runs[0].text = new_text
for run in para.runs[1:]:
run.text = ''
Complete Solution Pattern
from docx import Document
import json
import re
def fill_template(template_path, data_path, output_path):
"""Fill Word template handling all edge cases."""
# Load data
with open(data_path) as f:
data = json.load(f)
# Load template
doc = Document(template_path)
def replace_in_para(para):
text = para.text
pattern = r'\{\{([A-Z_]+)\}\}'
if not re.search(pattern, text):
return
new_text = text
for match in re.finditer(pattern, text):
key = match.group(1)
placeholder = match.group(0)
if key in data:
new_text = new_text.replace(placeholder, str(data[key]))
if new_text != text and para.runs:
para.runs[0].text = new_text
for run in para.runs[1:]:
run.text = ''
# Main document
for para in doc.paragraphs:
replace_in_para(para)
# Tables (with nesting)
def process_table(table):
for row in table.rows:
for cell in row.cells:
for para in cell.paragraphs:
replace_in_para(para)
for nested in cell.tables:
process_table(nested)
for table in doc.tables:
process_table(table)
# Headers/Footers
for section in doc.sections:
for para in section.header.paragraphs:
replace_in_para(para)
for para in section.footer.paragraphs:
replace_in_para(para)
doc.save(output_path)
# Usage
fill_template('template.docx', 'data.json', 'output.docx')
Common Pitfalls
- Forgetting headers/footers - They're not in
doc.paragraphs - Missing nested tables - Must recurse into
cell.tables - Split placeholders - Always work at paragraph level, not run level
- Losing formatting - Keep first run's formatting when rebuilding
- Conditional markers left behind - Remove
{{IF_...}}markers after processing
Related Skills
Agent-Reach
90.5kGive your AI agent eyes to see the entire internet. Read & search Twitter, Reddit, YouTube, GitHub, Bilibili, XiaoHongShu — one CLI, zero API fees.
headroom
74.4kCompress tool outputs, logs, files, and RAG chunks before they reach the LLM. 20% fewer tokens for coding agents, 60-95% fewer tokens for JSON, same answers. Library, proxy, MCP server.
Scrapling
85.6k🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl! Don't be shy, join here: https://discord.gg/EMgGbDceNQ and follow here for daily tips and tricks: https://x.com/Scrapling_dev
crawl4ai
84.8kOpen-source web crawler and scraper for LLMs and AI agents: any website into clean, LLM-ready Markdown. Run it yourself, or use Crawl4AI Cloud with one key.
Languages
Trust signals
From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.
