Memory prices tipped to fall as China starts flooding the market with DRAM and NAND chips

sanitation@lemmy.radio · 2 days ago

Memory prices tipped to fall as China starts flooding the market with DRAM and NAND chips

ImmersiveMatthew@sh.itjust.works · 20 hours ago

I just updated my setup from LMStudio to llama.cpp with the new QWEN 3.6 27B MTP model and I am getting 80-112 tokens/second, 90 average which is just shocking to me. I am on a 4090 with a context Window of 64k. It hardly use cloud AI anymore as I rarely need more than 64k if I ensure my first prompt is written like a design document. Multiple prompts are not great so I often just figure out where my initial prompt went wrong, adjust and try again in a fresh session. Way faster this way too. It has really worked out well for me as I am getting just as much done locally for free as I was with hundreds of dollar a month on cloud AI. I am still shocked and grateful it flowed this way.