Paper proposes two-level reinforcement learning for conversational agents
ToSCA proposes a hierarchical reinforcement-learning framework that separates strategic choices from token-level response generation. The authors report improved strategy selection and response quality against several baselines in daily and emotional-support conversations.