Notes on models for sequential data - RNNs, LSTMs, GRUs, and attention mechanism. Covers the pre-Transformer lineage of architectures for handling temporal dependencies, vanishing gradients, and memory in sequences.
Notes on models for sequential data - RNNs, LSTMs, GRUs, and attention mechanism. Covers the pre-Transformer lineage of architectures for handling temporal dependencies, vanishing gradients, and memory in sequences.
1 item with this tag.