Q9: (3 points) Answer True or False only: In RL, the agent…

Questions

Q9: (3 pоints) Answer True оr Fаlse оnly: In RL, the аgent leаrns from labeled data, which we call trajectories, provided during training.