Apple is considering offering an AI server built around "M8" series chips and has discussed incorporating Nvidia networking hardware, The Information reports. The system would be sold to outside customers, potentially bringing Apple back into a business it left behind when it discontinued Xserve in 2011. Apple is reportedly targeting companies that want to run AI models on their own equipment, with a particular focus on inference and generating responses from trained models.
Trust me, consumers aren’t the ones buying 5k€ displays, nor the nearly 20k€ maxed out Mac Studios.
Feels like opposite of “enterprise friendly priced hardware”
We don’t know what the pricing will be.
Something to consider: All those nVidia GPU servers have discrete GPUs, meaning they need to have both RAM and VRAM, and any time there’s need to transfer something between RAM and VRAM, that’s overhead. Apple runs everything in a shared pool of memory - which at present is significantly slower than the VRAM on those nVidia GPUs since it’s DDR rather than HBM, but if they do what they did with the Ultra line of chips and go even further, e.g glue together 8 chips instead of 2, they might make up a bit of the difference in memory bandwidth. Or they could add HBM to these chips I guess.
And a single nVidia DGX B300 with 2.1 TB VRAM is several hundred thousand. And a single one of those is not enough to run the biggest models. By offering less powerful chips and cheaper memory, they might theoretically be able to offer tons of memory and acceptable, though lower, performance for much less money. Allows security-conscious companies to run high-end LLMs locally. And since they deal in shared memory, you might be able to use CPU instead of GPU if that’s more efficient in some specific part of inference, without transferring things between different memory spaces.
Trust me, consumers aren’t the ones buying 5k€ displays, nor the nearly 20k€ maxed out Mac Studios.
We don’t know what the pricing will be.
Something to consider: All those nVidia GPU servers have discrete GPUs, meaning they need to have both RAM and VRAM, and any time there’s need to transfer something between RAM and VRAM, that’s overhead. Apple runs everything in a shared pool of memory - which at present is significantly slower than the VRAM on those nVidia GPUs since it’s DDR rather than HBM, but if they do what they did with the Ultra line of chips and go even further, e.g glue together 8 chips instead of 2, they might make up a bit of the difference in memory bandwidth. Or they could add HBM to these chips I guess.
And a single nVidia DGX B300 with 2.1 TB VRAM is several hundred thousand. And a single one of those is not enough to run the biggest models. By offering less powerful chips and cheaper memory, they might theoretically be able to offer tons of memory and acceptable, though lower, performance for much less money. Allows security-conscious companies to run high-end LLMs locally. And since they deal in shared memory, you might be able to use CPU instead of GPU if that’s more efficient in some specific part of inference, without transferring things between different memory spaces.