Reinforcement Learning. Reinforcement learning is a subfield of AI/statistics focused on exploring/understanding complicated environments and learning how to optimally acquire rewards. Examples are AlphaGo, clinical trials & A/B tests, and Atari game playing.
90.2K members. The best post we have captured here reached 7.6K upvotes, which works out to 84 per 1,000 members.
Open r/reinforcementlearning on Reddit ↗
The mix of post types among everything we have captured. Posting a link where almost nothing but images lands is wasted effort, and this is the fastest way to see it.
Flairs are the community's own categories, and the closest thing a subreddit has to telling you what it wants.
The biggest posts we have captured here. Links go to the original on Reddit.
Google copied our open-source code, removed our engineers’ names, and gave us zero credit one year after we beat them on their own benchmark
We beat Google Deepmind but got killed by a chinese lab
Trained a PPO agent to beat Lace in Hollow Knight: Silksong
IT'S LEARNING!
I literally build the jev architecture one year back and made it open-sourced
PPO Ping Pong
Why is RL fine-tuning on LLMs so easy and stable, compared to the RL we're all doing?
Andrew G. Barto and Richard S. Sutton named as recipients of the 2024 ACM A.M. Turing Award
Can RL redefine AI vision? My experiments with partial observation & Loss as a Reward
Built a custom robotic arm environment and trained an AI agent to control it
Within four times the member count either way, so the comparison is fair. Sorted by upside.
Answered from the posts ViralHunt holds from this community, not from Reddit's own about page.
ViralHunt tracks Reddit and eight other networks, records how every post grows, and puts what is taking off in front of your team in one list.
Start free →Every figure here is what ViralHunt has captured from r/reinforcementlearning, not everything the community has ever posted. We deliberately do not publish an average score: our sample is weighted toward a community's best posts, so an average would be flattering and wrong. Peak per 1,000 members is computed the same way for every community listed. See how we measure.