Complete, runnable reward functions for TRL's GRPOTrainer. Every function here follows the current TRL reward-function signature: it accepts completions plus any extra dataset columns as keyword arguments, and returns a list[float] the same length as completions. Base mod. Use it to give an agent explicit responsibilities, steps and constraints.
CUSTOM KNOWLEDGE FEED
#correctness
1 cardsThis feed is generated directly from exact card hashtags; there is no separate feed-content copy.
★ 0◌ 0