18 repositories · 9 under 3,000 stars
Language models, agents, embeddings, speech and the tools around them.
Fast semantic deduplication of text datasets — strips near-duplicate chat turns before topic modelling so clusters come out clean.
Splits text into real sentences in 85 languages using small CPU models, including text with no punctuation.
Scores how good a retrieval-based AI answer is, so a change to Grasppy can be shown to have helped or hurt.
Knowing which language a paragraph is in — before you spend a model call finding out.
The opposite architecture from the LiteLLM-plus-custom-code path — batteries actually included: FastAPI, RAG, plugin loader, WebSocket chat, all in one image.
One interface over several nearest-neighbour backends — the throwaway in-memory index for a single conversation.
Per-prediction feature attribution — which of tsfresh's 700 features the surviving model actually leans on, and whether that reason is embarrassing.
Hyperparameter search that samples the promising regions and prunes hopeless trials early — a week of sweeps back as an afternoon.
Entity-and-relationship graph over documents with local and global query modes — GraphRAG's structure at a fraction of the cost, stored in PostgreSQL.
Labelled data maps with proper label placement — a diagnostic for cluster naming and a shareable marketing image.
Names clusters at every level of a hierarchy using sub-cluster structure and sample documents rather than a bag of keywords.
Apple's interactive embedding map with automatic cluster labelling — recommended as a design study for Grasppy's overview screen, not as a dependency.
Self-hosted machine-translation API — removes the per-character meter from running a bilingual site.
Five families of topic model behind one scikit-learn-style API — the head-to-head against BERTopic in one afternoon.
The standard way to turn a long thread into named subtopics with a parent/child hierarchy and a timeline view.
Focused text-chunking library — semantic and recursive splitters that respect message boundaries, so subtopics come out clean.
Builds a knowledge graph from text, clusters it into nested communities, writes a summary for each level.
Distills sentence-transformers into static embeddings ~15x smaller and hundreds of times faster on plain CPU.
No repository in this area carries every tag you picked.
A dozen repositories, opened and checked. The licence read, the last release dated, and the ones that did not make it named with the reason. It is the half most lists leave out.
No tracking pixels. One click to leave. The archive stays free either way.
We use your address to send the edition and nothing else. Confirm by email, leave in one click. How we handle it.