Skip to content

feat(search): add SQLite FTS5 post-build search index script - #801

Open
jeevantelukula wants to merge 1 commit into
TexasInstruments:masterfrom
jeevantelukula:sqlite_fts5
Open

jeevantelukula wants to merge 1 commit into
TexasInstruments:masterfrom
jeevantelukula:sqlite_fts5

Conversation

@jeevantelukula

Copy link
Copy Markdown
Collaborator

Sphinx's searchindex.js stores only boolean term membership with no TF-IDF or field weighting, works good for manual browser search but gives poor ranking for programmatic (MCP/AI agent) search queries.

Add build_searchindex.py that runs after sphinx-build via html make target, parses HTML output, extracts the weighted fields (based on title, code_blocks, image_alt and body_text) and writes searchindex.db with a BM25-ranked FTS5 index. MCP server queries this DB to return accurately ranked doc pages to AI agents.

Also add beautifulsoup4 additionally to docker deps.

@StaticRocket

Copy link
Copy Markdown
Member

What are we actually going to use these databases for? Are they just going to be distributed alongside the documentation or is this expected to hook into the existing search function? If it's the latter, couldn't this be achieved with a custom scorer function.

@jeevantelukula

Copy link
Copy Markdown
Collaborator Author

What are we actually going to use these databases for? Are they just going to be distributed alongside the documentation or is this expected to hook into the existing search function? If it's the latter, couldn't this be achieved with a custom scorer function.

It is going to be distributed alongside the documentation. The existing search functionality which uses searchindex.js is good enough for manual search filtering.
I'm implementing an MCP sever for processor sdk docs, for that fts5 based database approach will give better scoring based pages to the agent querying the MCP.

Ref: https://github.com/jeevantelukula/ti-processor-sdk-linux-skills/tree/master/mcp

@StaticRocket

Copy link
Copy Markdown
Member

In that case I would prefer if the packaging / release preparation scripts specifically for ti.com be carried somewhere internally. I don't really want to include this in the public builds as it increases build time and results in artifacts too big to distribute here.

Sphinx's searchindex.js stores only boolean term membership with no
TF-IDF or field weighting, works good for manual browser search but
gives poor ranking for programmatic (MCP/AI agent) search queries.

Add build_searchindex.py that optionally runs after sphinx-build via
html make target, parses HTML output, extracts the weighted fields
(based on title, code_blocks, image_alt and body_text) and writes
searchindex.db with a BM25-ranked FTS5 index. MCP server queries
this DB to return accurately ranked doc pages to AI agents.

Also add beautifulsoup4 additionally to docker deps.

Signed-off-by: Telukula Jeevan Kumar Sahu <j-sahu@ti.com>
@jeevantelukula

Copy link
Copy Markdown
Collaborator Author

In that case I would prefer if the packaging / release preparation scripts specifically for ti.com be carried somewhere internally. I don't really want to include this in the public builds as it increases build time and results in artifacts too big to distribute here.

Good point. Although, the addtional size we are adding with this db is around 2-3MB for a device. Considering the number of devices we have, I have made this optional. So, github pages won't generate this by default and internally using BUILD_FTS_INDEX=1 flag we can generate the database for ti.com.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants