Can Transformers Really Do It All? On the Compatibility of Inductive Biases Across Tasks
Research explores optimizing transformer architectures for specific datasets to understand optimal task-specific inductive biases beyond current scaling methods.