Vibe Coding Discover

AI Frameworks

Official implementation of Hierarchical Context Merging: Better Long Context Understanding for Pre-trained LLMs (ICLR 2024).

★ 455 forksPython—alinlab

Official ICLR 2024 implementation of HOMER, a training-free hierarchical KV-cache merging method that extends pre-trained LLM context limits (e.g. Llama-2) with lower memory. Ships patched LlamaForCausalLM, plus passkey-retrieval and PG19 perplexity scripts.

Use Cases

Extending the effective context window of pre-trained LLMs without retrainingMemory-efficient inference on long documents via compact KV-cache mergingLong-context passkey retrieval benchmarkingLanguage modeling perplexity evaluation on PG19 long documentsCombining context merging with YaRN scaling for 32k+ contextsReproducing a published ICLR 2024 long-context research method

Built With

Language
Python
Frameworks
PyTorch · HuggingFace Transformers · FlashAttention-2 · Accelerate · sentencepiece · YaRN

Tags

long-context · kv-cache · context-extension · llm-inference · llama-2 · memory-efficiency · attention · transformers · pytorch · research-implementation · hierarchical-merging · passkey-retrieval · perplexity-evaluation · training-free · flash-attention · iclr-2024