But more memory is available so a larger model can be loaded. It depends on the use case which is better. Any chat or voice model will have better UX with nvidia but document or code generation will be better with apple.
it's not a binary thing. At some point Apple becomes too slow with very large models. If you can just run a model at 1 token per second and it takes 30 mins to process a long context, it's useless