Generative AI Is an Engineering Disaster. A shockingly inefficient trillion-dollar project.

a month ago by Some_Emo_Chick to c/technology

load all comments
Catoblepas 153 points a month ago

But it’s so good at programming if you already know how to program! Surely that’s worth burning the planet and crashing the world economy??

path: 0 24783377, hotness: undefined, score: 153, children: 19
Bonje 59 points a month ago

Actually still no

https://github.com/JustVugg/colibri

Everyone was desperate to be first because capitalism. But we are getting good models without the insane build out requirement. Which will be hilarious to leave the cunts holding the bag. Not that the planet is better for it in the end.

path: 0 24783377 24783726, hotness: undefined, score: 59, children: 16
ImgurRefugee114 32 points a month ago

~1 token per second (storage bound gen4 nvme)... Some of us have places to be.

Don't get me wrong. Its impressive that it can run at all, but honestly the usecase is exceedingly narrow. You'd have better results with a structured quantized gpu-only gemma or qwen workflow. Quality over quantity, rely on validation and a structured process: lots of cross-model review and iteration loops with spec and test driven dev. You could probably get a working alpha by the time colibri set up the environment.

path: 0 24783377 24783726 24784008, hotness: undefined, score: 32, children: 11
Asafum 7 points a month ago

Yeah I'm just beginning my local AI journey on a 5080, tried Qwen3.6 27b Q4 and was getting like 1tps because of the vram overflow. Ran it over night at it was still chewing on generating a prompt for a sub agent when I got up in the middle of the night until it simply ended in some kind of "fetch failure" lol. I think I gave it something too large to tackle, but either way 1tps is kinda garbage.

path: 0 24783377 24783726 24784008 24784351, hotness: undefined, score: 7, children: 10
worldclasspun 4 points a month ago

Is that quantized? 4 bit Qwen 3.6 can get 22tps on a 1060.

path: 0 24783377 24783726 24784008 24784351 24784888, hotness: undefined, score: 4, children: 4
Asafum 3 points a month ago

It's the q4 quantization, but it requires 20+GB vram and my 5080 only has 16

path: 0 24783377 24783726 24784008 24784351 24784888 24786374, hotness: undefined, score: 3, children: 3
Damage 3 points a month ago

My framework 13 with shared RAM runs qwen quite well

path: 0 24783377 24783726 24784008 24784351 24784888 24786374 24791732, hotness: undefined, score: 3, children: 0
worldclasspun 1 point 10 days ago

What are you using to run the model? Llama.cpp will automatically split the model between your system ram and graphics card's vram.

Qwen 3.6 is a mixture of experts model with only 3B parameters active at a time. Even without quantization your card could easily run that.

path: 0 24783377 24783726 24784008 24784351 24784888 24786374 25239254, hotness: undefined, score: 1, children: 1
technology
technology

@lemmy.world

login for more options
87402
21272
15822

This is a most excellent place for technology news and articles.

Our Rules


  1. Follow the lemmy.world rules.
  2. Only tech related news or articles.
  3. Be excellent to each other!
  4. Mod approved content bots can post up to 10 articles per day.
  5. Threads asking for personal tech support may be deleted.
  6. Politics threads may be removed.
  7. No memes allowed as posts, OK to post as comments.
  8. Only approved bots from the list below, this includes using AI responses and summaries. To ask if your bot can be added please contact a mod.
  9. Check for duplicates before posting, duplicates may be removed
  10. Accounts 7 days and younger will have their posts automatically removed.

Approved Bots


go to feed...