a lightweight hybrid reasoning MoE model with 7.9B total parameters and only 1.3B activated parameters per token. It is designed to deliver strong reasoning and agentic capabilities under a small inference compute footprint, making advanced model capabilities more accessible for local and resource-constrained deployment.

  • fozid@lem.radiantfig.fyi
    link
    fedilink
    English
    arrow-up
    2
    ·
    14 days ago

    I haven’t tried it yet, but I tested one of their previous ones with CPU interference on an Intel n97, and it was one of the best in terms of t/s performance and also prompt response quality on the metrics I tested against. When this run on llama.cpp I will try and give it a test.

  • leanleft@lemmy.mlOP
    link
    fedilink
    English
    arrow-up
    1
    ·
    14 days ago

    personally,
    i did one test of ling3-flash and gemma-4-31B side-by-side .
    ling3 understood me and had a fantastic answer. gemma must have misunderstood what i was saying… it wrote a long story-like paragraph that didnt answer my question. but it kinda had 1 bit of insight.
    so: (maybe…) don’t sleep on ling models !