Turn your own notes into practice that sticks
Upload study material you already own. QuestionMill writes practice questions from it, grades your answers, and brings each one back on a spaced-repetition schedule.
Works with PDF, PPTX, DOCX, TXT and Markdown, plus images (PNG, JPG, GIF, WebP) — photos of handwritten notes work too.
- Upload material you already own — lecture notes, textbook excerpts, a revision guide.
- Answer practice questions generated from that material.
- See what you got wrong, ask follow-up questions grounded in the same document, and review on a schedule.
Create a free account in seconds · Pricing · Sign in
For self-directed learners aged 13 and over — university students, certification candidates, and anyone studying on their own.
Your material stays private to your account. We do not publish it, share it, or use it to train models.
See it work
One paragraph of your own notes in, one question out — with the passage it came from.
Your material Lecture notes · gradient descent
The learning rate is the single most important hyperparameter. If it is too large, each step overshoots the minimum and the loss can oscillate or diverge. If it is too small, convergence is slow and training may stall in a flat region before reaching a good solution.
What you practise
In gradient descent, what happens when the learning rate is set too large?
- Each step overshoots the minimum, so the loss can oscillate or diverge.
- Convergence becomes slow and training may stall in a flat region before reaching a good solution.
- The gradient noise helps the optimiser escape shallow local minima.
- Oscillation across steep ravines is damped because past gradients are averaged.
- The gradient of the loss with respect to the parameters is no longer subtracted from the parameters.
Why — The material states that the learning rate is the single most important hyperparameter and that if it is too large, each step overshoots the minimum and the loss can oscillate or diverge. Option B describes the opposite problem — a learning rate that is too small, which makes convergence slow and can cause training to stall in a flat region. …
Where it comes from — The learning rate is the single most important hyperparameter. If it is too large, each step overshoots the minimum and the loss can oscillate or diverge.