In this paper we discuss improving yield in semiconductor manufacturing using reinforcement learning (RL) to tune a dispatching rule parameter to increase the number of lots that process on high-yield equipment. We consider a dispatching rule with a parameter that controls whether or not the rule allows a lot to process on a lower-yield equipment or waits to allow the lot to possibly process later on a high-yield equipment. In a factory such a parameter would be set periodically, e.g., once a week, but RL allows the parameter to be updated frequently, leading to better factory performance. We also consider a novel measure of on-time delivery where the goal is to have 95% on-time delivery in a set of time intervals, e.g., shifts. We show how a trained RL agent using a graph neural network outperforms the baseline by maintaining on-time delivery while processing significantly more lots on high-yield equipment.