Diagram of an actor-critic architecture showing input state features to a shared MLP, which branches into separate ACTOR and CRITIC heads for portfolio weights and expected future reward, respectively. PPO algorithm details noted below.

keyboard_arrow_up