Reinforcement-Learning Policy Evaluator
Evaluate a fixed finite-state policy by iterating its reward vector and policy-induced transition matrix.
Description
Evaluate a fixed finite-state policy by iterating rewards and its transition matrix.
Reinforcement-Learning Policy Evaluator: Evaluate a fixed finite-state policy by iterating rewards and its transition matrix.
When to use Reinforcement-Learning Policy Evaluator
Use this reinforcement-learning update to study policy evaluation or gradient-based learning in a fully specified environment with explicit states, actions, rewards, discounting, and sampling behavior.
- State rewards
- Required list input.
- Transition matrix
- Required list input.
- Discount factor
- Required number input.
- Iterations
- Required integer input.
How Reinforcement-Learning Policy Evaluator works
Evaluate a fixed finite-state policy by iterating rewards and its transition matrix. The tool evaluates the supplied inputs together and returns the named outputs below; it does not infer omitted operating conditions or change the units shown.1
- State values
- The resulting state values returned as a list.
Limitations and assumptions
- Estimates can be biased or high-variance and depend on exploration, bootstrapping, function approximation, rollout length, advantage normalization, off-policy corrections, seeds, and environment stationarity. Training return does not establish safe deployment.
- Use finite inputs in the displayed units and preserve more precision than the final presentation requires. Independently verify safety-critical, financial, compliance, or production decisions.
Alternative or Complementary approaches
Test tabular cases, run multiple seeds, report confidence intervals and sample cost, evaluate off-policy and under perturbations, and apply domain safety constraints.
References
-
Reinforcement learning — Wikipedia contributors
Similar or alternative tools
- Policy-Gradient Training Step
Apply one categorical REINFORCE policy-gradient update from an action, return, and baseline.
- A3C Rollout Update Calculator
Calculate discounted returns, advantages, and A3C rollout losses.
- Finite-State Markov Chain Monte Carlo Sampler
Sample a finite-state Markov chain from a caller-supplied row-stochastic transition matrix and initial state. This samples the chain, not an arbitrary target distribution.