Hard Attention to the Task (HAT)
Learns a hard attention mask for each task that selectively gates hidden units. During training on new tasks, a gradient compensation mechanism prevents updates to units that are important for previous tasks. This provides near-zero forgetting with minimal capacity overhead.