Use AG2 to Tune ChatGPT - AG2

Use AG2 to Tune ChatGPT

AG2 offers a cost-effective hyperparameter optimization technique EcoOptiGen for tuning Large Language Models. The study finds that tuning hyperparameters can significantly improve the utility of LLMs. Please find documentation about this feature here.

In this notebook, we tune OpenAI ChatGPT (both GPT-3.5 and GPT-4) models for math problem solving. We use the MATH benchmark for measuring mathematical problem solving on competition math problems with chain-of-thought style reasoning.

Requirements

AG2 requires Python>=3.9. To run this notebook example, please install with the [blendsearch] option:

pip install "ag2[blendsearch]<0.2"

AG2 has provided an API for hyperparameter optimization of OpenAI ChatGPT models: autogen.ChatCompletion.tune and to make a request with the tuned config: autogen.ChatCompletion.create. First, we import autogen:

import datasets

import autogen
from autogen.math_utils import eval_math_responses

Set your API Endpoint

The config_list_openai_aoai function tries to create a list of Azure OpenAI endpoints and OpenAI endpoints. It assumes the api keys and api bases are stored in the corresponding environment variables or local txt files:

It’s OK to have only the OpenAI API key, or only the Azure OpenAI API key + base.

config_list = autogen.config_list_openai_aoai()

Load dataset

We load the competition_math dataset. The dataset contains 201 “Level 2” Algebra examples. We use a random sample of 20 examples for tuning the generation hyperparameters and the remaining for evaluation.

seed = 41
data = datasets.load_dataset("competition_math")
train_data = data["train"].shuffle(seed=seed)
test_data = data["test"].shuffle(seed=seed)
n_tune_data = 20
tune_data = [
    {
        "problem": train_data[x]["problem"],
        "solution": train_data[x]["solution"],
    }
    for x in range(len(train_data))
    if train_data[x]["level"] == "Level 2" and train_data[x]["type"] == "Algebra"
][:n_tune_data]
test_data = [
    {
        "problem": test_data[x]["problem"],
        "solution": test_data[x]["solution"],
    }
    for x in range(len(test_data))
    if test_data[x]["level"] == "Level 2" and test_data[x]["type"] == "Algebra"
]
print(len(tune_data), len(test_data))

Check a tuning example:

print(tune_data[1]["problem"])

Here is one example of the canonical solution:

print(tune_data[1]["solution"])

Define Success Metric

Before we start tuning, we must define the success metric we want to optimize. For each math task, we use voting to select a response with the most common answers out of all the generated responses. We consider the task successfully solved if it has an equivalent answer to the canonical solution. Then we can optimize the mean success rate of a collection of tasks.

Use the tuning data to find a good configuration

For (local) reproducibility and cost efficiency, we cache responses from OpenAI with a controllable seed.

autogen.ChatCompletion.set_cache(seed)

This will create a disk cache in “.cache/{seed}”. You can change cache_path_root from “.cache” to a different path in set_cache(). The cache for different seeds are stored separately.

Perform tuning

The tuning will take a while to finish, depending on the optimization budget. The tuning will be performed under the specified optimization budgets.

Users can specify tuning data, optimization metric, optimization mode, evaluation function, search spaces etc. The default search space is:

default_search_space = {
    "model": tune.choice([
        "gpt-3.5-turbo",
        "gpt-4",
    ]),
    "temperature_or_top_p": tune.choice(
        [
            {"temperature": tune.uniform(0, 2)},
            {"top_p": tune.uniform(0, 1)},
        ]
    ),
    "max_tokens": tune.lograndint(50, 1000),
    "n": tune.randint(1, 100),
    "prompt": "{prompt}",
}

The default search space can be overridden by users’ input. For example, the following code specifies a fixed prompt template. The default search space will be used for hyperparameters that don’t appear in users’ input.

prompts = [
    "{problem} Solve the problem carefully. Simplify your answer as much as possible. Put the final answer in \boxed{{}}."
]
config, analysis = autogen.ChatCompletion.tune(
    data=tune_data,  # the data for tuning
    metric="success_vote",  # the metric to optimize
    mode="max",  # the optimization mode
    eval_func=eval_math_responses,  # the evaluation function to return the success metrics
    inference_budget=0.02,  # the inference budget (dollar per instance)
    optimization_budget=1,  # the optimization budget (dollar in total)
    num_samples=20,
    model="gpt-3.5-turbo",  # comment to tune both gpt-3.5-turbo and gpt-4
    prompt=prompts,  # the prompt templates to choose from
    config_list=config_list,  # the endpoint list
    allow_format_str_template=True,  # whether to allow format string template
)

Output tuning results

After the tuning, we can print out the config and the result found by AG2, which uses flaml for tuning.

print("optimized config", config)
print("best result on tuning data", analysis.best_result)

Make a request with the tuned config

We can apply the tuned config on the request for an example task:

response = autogen.ChatCompletion.create(context=tune_data[1], config_list=config_list, **config)
metric_results = eval_math_responses(autogen.ChatCompletion.extract_text(response), **tune_data[1])
print("response on an example data instance:", response)
print("metric_results on the example data instance:", metric_results)

Evaluate the success rate on the test data

You can use autogen.ChatCompletion.test to evaluate the performance of an entire dataset with the tuned config. The following code will take a while (30 mins to 1 hour) to evaluate all the test data instances if uncommented and run. It will cost roughly $3.

# result = autogen.ChatCompletion.test(test_data, logging_level=logging.INFO, config_list=config_list, **config)
# print("performance on test data with the tuned config:", result)

What about the default, untuned gpt-4 config (with the same prompt as the tuned config)? We can evaluate it and compare:

# default_config = {"model": 'gpt-4', "prompt": prompts[0], "allow_format_str_template": True}
# default_result = autogen.ChatCompletion.test(test_data, config_list=config_list, **default_config)
# print("performance on test data from gpt-4 with a default config:", default_result)

The default use of GPT-4 has a much lower accuracy. Note that the default config has a lower inference cost. What if we heuristically increase the number of responses n?

# config_n2 = {"model": 'gpt-4', "prompt": prompts[0], "n": 2, "allow_format_str_template": True}
# result_n2 = autogen.ChatCompletion.test(test_data, config_list=config_list, **config_n2)
# print("performance on test data from gpt-4 with a default config and n=2:", result_n2)

The inference cost is doubled and matches the tuned config. But the success rate doesn’t improve much. What if we further increase the number of responses n to 5?

# config_n5 = {"model": 'gpt-4', "prompt": prompts[0], "n": 5, "allow_format_str_template": True}
# result_n5 = autogen.ChatCompletion.test(test_data, config_list=config_list, **config_n5)
# print("performance on test data from gpt-4 with a default config and n=5:", result_n5)

We find that the ‘success_vote’ metric is increased at the cost of exceeding the inference budget. But the tuned configuration has both higher ‘success_vote’ (91% vs. 87%) and lower average inference cost ($0.015 vs. $0.037 per instance).

A developer could use AG2 to tune the configuration to satisfy the target inference budget while maximizing the value out of it.