Neutrino-1 8B

(fermionresearch.com)

78 points | by handfuloflight 4 hours ago

11 comments

  • kamranjon 3 hours ago
    There's a really interesting trend of labs using "proprietary" methods to convert existing models to a compressed ternary format.

    PrismML actually targeted the same Qwen 8b model and got it down to 1.75gb here: https://prismml.com/news/ternary-bonsai

    I wonder how proprietary it all is though, since the BitNet b1.58 paper has been out for a couple years now: https://arxiv.org/abs/2402.17764

    From the wikipedia on 1.58 bit llms: "BitNet derives its performance from being trained natively in 1.58 bit instead of being quantized from a full-precision model after training. Still, training is an expensive process, and it would be desirable to be able to somehow convert an existing model to 1.58 bits. In 2024, HuggingFace reported a way to gradually ramp up the 1.58-bit quantization in fine-tuning an existing model down to 1.58 bits."

    The section from huggingface is here: https://huggingface.co/blog/1_58_llm_extreme_quantization#fi...

    I just wonder how many of these labs are basically following the huggingface recipe here and possibly tweaking it and releasing models without huge training costs.

    • Twirrim 3 hours ago
      Independent testing of prismml suggest quite a capability drop off outside of their cherry picked benchmarks. I'll be curious to see what this model achieves though.
      • kamranjon 3 hours ago
        Unfortunately Fermion Research appears to entirely AI generate all of their content here, even for the research section: https://www.fermionresearch.com/research/neutrino-8b/

        "Neutrino-1 8B was trained natively in its shipping format. There is no full-precision product model that was rounded afterward: the ternary representation is the medium the weights learned in, and the training methods that hold this quality at this depth are the lab’s unpublished work. The findings below are the part that travels."

        This statement seems misleading at best.

        Both the model page and the release page are basically unintelligible - I don't have a ton of faith in the work here, at least PrismML write coherent releases for their models.

        Edit: Another beautiful piece of prose here, I almost wonder if they used the 8b model to generate the content for this release...

        "Across the 6.95B coded weights, 62.63% sit at zero and the remainder splits 18.68% plus to 18.69% minus: sign-balanced to a hundredth of a point with no constraint asking for it."

        • LtdJorge 51 minutes ago
          I guess it's saying how many of the weights are -1, 0 or +1.
      • dofm 1 hour ago
        I really had high hopes for the larger Ternary Bonsai and it feels like there is scope to improve, but I get the sense (albeit a naïve, probably not fully informed sense) that improvement can perhaps only come by training directly into ternary.
        • kamranjon 44 minutes ago
          I’ve actually been really impressed with the 27b model they recently released - amazing performance approaching 40 tok/s on m4 max and I didn’t run into any quality issues in the small set of tasks I tried. Haven’t gone full coding with it yet but suspect it’s better than say a 9b or 12b model.
        • avadodin 42 minutes ago
          All you need is Ternary Aware Training and for AI researchers to come up with a backronym for TIT.
  • moinism 3 hours ago
    The content on that page is too AI-generated to make sense to me; I don't understand what the model is for.
    • naruhodo 50 minutes ago
      I have the same question.

      I get that it’s designed to run on a CPU, big GPU or MacBook (although the way that was phrased confused me at first).

      I’m struggling with what a “decoder-only” model is good for.

    • SwellJoe 2 hours ago
      AI doesn't want anything, so it doesn't care whether it conveys meaning in its writing. And, apparently the developers of this project also don't care whether it conveys meaning. They just assume we'll wade through the slop? I dunno.
  • secult 1 hour ago
    There is not a single person mentioned on the website, github created 3 days ago, no real contact, everything hidden. Completely anonymous. Domain owner hidden.
  • Havoc 1 hour ago
    Can’t say I’m a fan of containers for this. A big chunk of local LLM gains come (imo) from the open modular nature of llama.cpp and friends. Easy to modify. Easy to experiment.

    Containers are the proprietary binary blob in hardware world equivalent

    • embedding-shape 11 minutes ago
      What? How are those even related with each other? You can just as easy modify and experiment with llama.cpp in a container as outside of it, they really shouldn't impact one another. Containers don't suddenly make llama.cpp less "open modular" somehow, and I'm not sure how you'd arrive as such conclusion.
  • yborg 3 hours ago
    Largely outperformed by Ternary-Bonsai-8B by their own chart, doesn't seem clear what their special sauce is here.
  • NetOpWibby 2 hours ago
    There’s a new announcement every other day wrt models. How do y’all keep track of them all same know what’s decent? Good grief!

    And if it’s decent today, it’s shit in eight months! I tool hop as much as the next dev but this is a bit much.

    • weikju 2 hours ago
      I just tune out. It’s not worth knowing, following every development in the field. If something works now it will probably work in 8 months even if it’s no longer the new hype thing. Who cares.

      Not using any of it is also a valid option though it doesn’t satisfy your FOMO. But nothing ever will.

    • badatnames 48 minutes ago
      There's no need to, the 50 foot view is simply that many alternatives exist and they mostly fall into 3 meaningful weight classes with comparable performance among each class's members: too expensive to use indiscriminately, too big to run at home, and too small for complex work. As for names and faces in between, the overarching conclusion is that we're rapidly approaching commodity status and those don't really matter much
  • madhu_ghalame 2 hours ago
    The real strength of an 8B model is efficiency. It will be interesting to see the balance between performance and inference cost.
  • Alien1Being 1 hour ago
    AI slop site with AI slop research...

    Blog populated with incoherent PR material generated by Yet Another AI.

    Sigh...

    • zoom6628 40 minutes ago
      Can Dang implement a slop rating on submitted pages? Not a block but at least a % likelihood of AI slop content and that could also be tied with a BS rating as well.

      Could use AI for both which seems hilariously appropriate.

  • drbscl 1 hour ago
    Slop article, slop site... slop model?
  • runtime_lens 2 hours ago
    [flagged]