Method: compared “stars today” from GitHub Trending daily snapshots (Sep 26, 22:54 UTC) against the previous day’s snapshot (Sep 26, 00:49 UTC) and ranked by recent gain. Repos above 40k total stars or with slowing momentum were excluded (paperclipai/paperclip at 86k, rohitg00/ai-engineering-from-scratch at 58k, and dream-num/univer, which slowed from +1,050 to +845). Weekend trending was thin, so only two repos remain.
vectorize-io/hindsight — +2,152 ⭐ today
URL: https://github.com/vectorize-io/hindsight
Total: ~31–32k ⭐ (previous day +1,653 → today +2,152, about 30% acceleration) · Python · MIT
Hindsight is an agent memory system designed to make agents learn over time rather than simply replay conversation history. It is built around three operations:
- Retain — uses an LLM to extract facts, entities, relationships and temporal information and stores them.
- Recall — retrieves memories by running four strategies in parallel: semantic search, keyword matching, entity/temporal/causal graph traversal and time-based filtering.
- Reflect — forms new connections among stored memories to build a synthesized understanding.
Memories are organized into world facts, experiences, observations (deduplicated beliefs strengthened by evidence) and mental models (standing answers about a memory bank), a structure closer to human memory than a plain vector database. It supports 25+ LLM providers, 60+ integrations including LangChain, CrewAI, Claude Code and Cursor, an MCP server and a two-line LLM wrapper, and deploys via Docker, Helm or as an embedded library; the project reports beating RAG and knowledge-graph approaches on LongMemEval. After reaching the top of the weekly charts in mid-September, its daily gains are growing again this weekend — a clear sign that keeping context across long-running agent sessions has become a core industry problem.
Practical use: Give customer-support or internal coding agents per-user, isolated long-term memory so they carry over preferences and past decisions without repeated explanations; PII redaction and PostgreSQL/Oracle support make enterprise evaluation easier.
Tags: #AgentMemory #LongTermMemory #KnowledgeGraph #LongMemEval #MCP #ContextEngineering
NVIDIA/Model-Optimizer — +354 ⭐ today
URL: https://github.com/NVIDIA/Model-Optimizer
Total: ~4.6–4.8k ⭐ (previous day +359 → today +354, holding steady two days running) · Python · Apache-2.0
NVIDIA Model Optimizer (ModelOpt) is a unified library of state-of-the-art model compression techniques. It takes Hugging Face, PyTorch or ONNX models, optimizes them, and exports checkpoints ready for deployment frameworks. Supported techniques:
- Post-training quantization (PTQ): 2–4x compression without retraining
- Quantization-aware training (QAT): recovers accuracy of quantized models through training
- Pruning and sparsity
- Knowledge distillation
- Draft-module training for speculative decoding
- Neural architecture search (NAS)
Outputs deploy to TensorRT-LLM, TensorRT, vLLM, SGLang and Diffusers models, and training-based optimizations integrate with Megatron-Bridge, Megatron-LM and Hugging Face Accelerate. Renamed from TensorRT Model Optimizer and open-sourced in January 2025, it recently added AutoQuantize for fast automatic mixed-precision assignment (August) and published W4A4 NVFP4 plus quantization-aware distillation results for Qwen3.6-35B with 1.30x vLLM throughput (September). A small repo steadily adding around 350 stars a day for two days signals that demand for cheaper inference is converging on NVIDIA’s official toolchain.
Practical use: Quantize open-weight LLMs to NVFP4/FP8 on a customer’s on-prem GPUs to raise vLLM or SGLang serving throughput and cut GPU count in a cost-optimization PoC.
Tags: #Quantization #NVFP4 #KnowledgeDistillation #SpeculativeDecoding #TensorRTLLM #InferenceOptimization
Leave a comment