C.W.K.
Stream
Quiz · 5 questions

📝 Sequences, RNNs & Attention

From SimpleRNN to LSTM to Multi-Head Attention — and why we got here

Level 0Level 0
0 XP0/78 lessons0/17 achievements
0/100 XP to next level100 XP to go0% complete

Quiz

01Why does a stacked LSTM need return_sequences=True on all but the last layer?
02Why does GRU train faster than LSTM with similar accuracy?
03What's the architectural reason Transformers outperform LSTMs on long sequences?
04What does merge_mode='concat' do in a Bidirectional LSTM wrapper?
05Why include TextVectorization inside the model rather than preprocessing text separately?
Spotted a bug or have feedback on this page?Report an Issue

Comments 0

🔔 Reply notifications (sign in)
Sign inPlease sign in to comment.

No comments yet — be the first.