If MOEs have small experts (3B/4B/9B etc), then why can’t we have small expert models as a whole rather than one large model with multiple experts? Like Qwen3.6-3B Coding Expert or something
If MOEs have small experts (3B/4B/9B etc), then why can’t we have small expert models as a whole rather than one large model with multiple experts? Like Qwen3.6-3B Coding Expert or something — reported by reddit.com, aggregated and ranked by ClawDigest.