Can you self-host AI at parity with chatgpt?

wuphysics87@lemmy.ml · 4 months ago

Can you self-host AI at parity with chatgpt?

TootGuitar@sh.itjust.works · edit-2 4 months ago

It depends on what you mean by “relative responsiveness”, but you can absolutely get ~4 tokens/sec of performance on R1 671b (Q4 quantized) from a system costing a fraction of the number you quote.

Xanza@lemm.ee · 4 months ago

This is the point everyone downvoting me seems to be missing. OP wanted something comparable to the responsiveness of chat.chatgpt.com… Which is simply not possible without insane hardware. Like sure, if you don’t care about token generation you can install an LLM on incredibly underpowered hardware and it technically works, but that’s not at all what OP was asking for. They wanted a comparable experience. Which requires a lot of money.

TootGuitar@sh.itjust.works · 4 months ago

Yeah I definitely get your point (and I didn’t downvote you, for the record). But I will note that ChatGPT generates text way faster than most people can read, and 4 tokens/second, while perhaps slower than reading speed for some people, is not that bad in my experience.