TLDR: It’s here, and it will suck for you FOR NOW. IT WILL GET BETTER. So, we have managed to have Kimi-K3 running on Blablador. As far as I can tell, this is the only lab in the world which managed to have it running on A100 GPUs in any production-like deployment. Sure, there are people claiming that they can run it on a Mac Mini. It’s like saying that I have seen, dunno, Taylor Swift, when all I saw was a tattoo of her face on someone’s leg. Sure, fam. Go on. We are running the full thing. No quantizations, no optimizations, no tricks. Just as Moonshot AI recommended. Except we are using hardware they never expected it to run. For that, we did a number of crazy things here. I will document this eventually - it deserves at least a blog post, better a paper. Anyway, what I wanted to say is: hold your horses. This is EXPERIMENTAL. This is SLOW. Kimi has always been extremely verbose. It might take HOURS to get a reply. The replies are good, though. What I mean is: don’t point your agents to it and complain later that the thing takes forever. It does. And I will be killing the deployment multiple times before my vacations are over. It won’t be there as available as Qwen is. That being said, and I hope I tone down the expectations, the model is GOOD. It can go for hours on tasks, slowly for now. And, as I said, I will be killing the deployment to run experiments. Don’t expect stability for now. As it gets better, it will become part of our permanent roster of models. I also want it. Keep barking, Alex, the one who barks for random dogs and they bark back
participants (1)
-
Strube, Alexandre