Global Tech News Technology. People. A more open tomorrow.
AI Inference Split Raises Tokens per Joule by Up to 56% in GPU Proxy Chips

AI Inference Split Raises Tokens per Joule by Up to 56% in GPU Proxy

Researchers at Harvard University and Nvidia have proposed a way to divide long-context artificial-intelligence inference between two types of processor. In an adjusted test using eight Nvidia B200 graphics processors as a hardware proxy, their design produced 31…

Latest Technology

View all latest news