Gradient Episodic Memory (GEM)
Stores a small episodic memory of past examples and uses them to constrain gradient updates during training on new tasks. GEM projects the gradient onto a feasible region that does not increase the loss on previous tasks, providing formal guarantees against forgetting.