Skip to content

Expose APIs from RL4LM to Consume 2 Models, 1 Scenario and 1 Morality Score to align the schema generator LM #8

Description

@logisticloon

The purpose of this subtask is to ensure Alignment of LM_SchemaGenerator

The two models are LM_SchemaGenerator and LM_Task.

The API should expect path of the two models (perhaps checkpoints and Another API callback function for LM_Task) and a scenario as input parameters.

Ensure a fork from RL4LM and not RL4F.

Subtasks:

  • Extend the metric function to incorporate MSE ( Reuse your last implementation )
  • PPO training setup for LM_SchemaGenerator.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

No labels
No labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions