NHacker Next
  • new
  • past
  • show
  • ask
  • show
  • jobs
  • submit
Run QWEN3.8 27B on 16gb Nvidia GPUs (github.com)
Pragmata 3 hours ago [-]
I've been using it for the past few days, and it runs really well!

I usually get 7 token/s using llama or lm studio, but this inference recipe runs at a smooth 80 tokens per second.

Genuinely very usable, and fully local!

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact
Rendered at 22:29:39 GMT+0000 (Coordinated Universal Time) with Vercel.