We have hosted the application colibri in order to run this application in our online workstations with Wine or directly.


Quick description about colibri:

Colibri is a compact inference engine designed to run the 744-billion-parameter GLM-5.2 mixture-of-experts model on consumer hardware. It keeps the dense portion of the quantized model in memory while streaming routed experts from a large disk-based store as they are needed. The runtime is implemented in pure C, requires no Python or BLAS during inference, and can operate without a GPU. Compressed attention caches, expert caching, optional hot tiers, and speculative decoding reduce memory pressure and improve repeated use. A planning tool calculates safe disk, RAM, and VRAM placement before loading the model, while a diagnostic command checks system readiness. Colibri includes terminal chat, an OpenAI-compatible text API, and a browser client, but disk-bound generation can be slow on cold caches.

Features:
  • GLM-5.2 inference on consumer hardware
  • Disk-streamed mixture-of-experts architecture
  • Dependency-free pure C inference runtime
  • Compressed KV cache and expert caching
  • Automatic RAM and VRAM placement planning
  • Terminal chat and OpenAI-compatible API


Programming Language: C.
Categories:
AI Models

Page navigation:

©2024. Winfy. All Rights Reserved.

By OD Group OU – Registry code: 1609791 -VAT number: EE102345621.