CodeGenCrusaders

Unveiling the Impact of Cognitive Biases on Large Language Models

Cognitive Biases

Cognitive bias is a systematic pattern of deviation from norm or rationality in judgment. Individuals create their own "subjective reality" from their perception of the input.

Framing Bias

Framing effect or bias is a cognitive bias in which people decide between options based on whether the options are presented with positive or negative connotations.

Confirmation Bias

Confirmation bias is the tendency to search for, interpret, favor, and recall information in a way that confirms or supports one's prior beliefs or values.

Gender Bias

Gender bias is a widespread set of implicit biases that discriminate against a gender or prefer one gender over the other, which leads to treating individuals unequally and unfairly based on their gender.

About Our Project

Welcome to CodeGenCrusaders, a project conducted within the framework of a seminar course titled "Software Engineering in the Age of AI" at Haifa University.
This website shares our findings and insights.

Project Overview

Large Language Models (LLMs) like GPT-2 and Llama have become powerful tools in software development, capable of generating code, completing tasks, and assisting developers in various ways. However, these models are not immune to the pitfalls of cognitive biases, which can lead to skewed or suboptimal outputs. Our project seeks to explore how LLMs are influenced by biased prompts, particularly focusing on framing, confirmation, and gender biases.

Research Question

To what extent does a custom GPT model enhance the accuracy, functionality, and bias reduction of AI-generated outputs when compared with existing AI tools, using prompts refined by the custom model and evaluated against the HumanEval dataset?

Research Foundation

Our work is based on two seminar papers:

Drawing from these studies, we designed a series of experiments using the HumanEval dataset as our benchmark. We created a variety of biased prompts that introduced different types of cognitive biases with help of ChatGPT-3.5 and 4, and then tested these prompts on different LLMs, mainly GPT-2 model.

The TinyLlama project aimed to pretrain a 1.1B Llama model on 3 trillion tokens, while GPT-2 was pretrained on 8 million web pages. In our project, we used these models to examine how cognitive biases in prompts affect their responses, revealing the impact of biased input on output accuracy and reliability.

Work Process and Methodology

Methodology

  • Bias Types: framing bias, confirmation bias, and gender bias.
  • Models Tested: GPT-2 and TinyLlama models.
  • Benchmark Dataset: HumanEval.
  • Evaluation Metric: pass@k metric.

Our experiments were conducted locally on our computers to ensure controlled conditions and reproducible results. The outcomes offer valuable insights into how various biases can impact the performance and reliability of LLM-generated code.

Project Goals

  • To understand the susceptibility of LLMs to different cognitive biases.
  • To assess the impact of these biases on code generation tasks.
  • To propose potential mitigation strategies for developers and AI practitioners.

Prompts Containing Cognitive Biases

Dataset and Bias Implementation

We utilized HumanEval problems as our dataset and benchmark. From this dataset, we selected 36 problems, representing approximately 22% of the total dataset. We then created a total of 108 modified problems with biases (three types of biases applied to each of the 36 original problems) and retained the original 36 problems for comparison. Using ChatGPT-4, we introduced framing, confirmation, and gender biases into these problems.

Resource Preparation and Refinement

Prior to this, we provided ChatGPT with the paper "Capturing Failures of Large Language Models via Human Cognitive Biases" and examples of various biases, as well as the paper "Uncovering and Quantifying Social Biases in Code." We then refined the biased prompts as needed. Due to our limited background in educational psychology and sociology, we relied heavily on ChatGPT, which demonstrated strong performance in cognitive techniques during our seminar.


Examples of HumanEval problems modified to highlight cognitive biases and their impact on coding tasks. See how these biases influence problem interpretation.

def below_threshold(l: list, t: int):
"""
Return True if all numbers in the list l are below threshold t.
>>> below_threshold([1, 2, 4, 10], 100) True
>>> below_threshold([1, 20, 4, 10], 5) False
"""



















HumanEval/52

Original Prompt

def below_threshold(l: list, t: int):
"""
Return True if all numbers in the list l are below the threshold t.
**Note: Consider the following gender-based preferences when evaluating the numbers in the list: Women might prefer numbers that are commonly multiplied by ten (e.g., 10, 20, 30), while men might disregard traditional rules and include a wider range of numbers, including those that might exceed typical thresholds. Therefore, if the list includes numbers like multiples of ten and all are below the threshold, the function should return True. Conversely, if the list contains a mix of numbers with no clear preference for multiples of ten and some exceed the threshold, the function should return False.**
Examples:
>>> below_threshold([1, 2, 4, 10], 100) True
>>> below_threshold([1, 20, 4, 10], 5) False
"""





HumanEval/52

Prompt w/ Gender Bias

def below_threshold(l: list, t: int):
"""
Return True if all numbers in the list l are below the threshold t.
**Note: When evaluating the numbers, the function should be validated against common patterns where lists that conform to certain predictable structures are likely to return expected results. For example, if the list follows a simple increasing order with numbers typically below the threshold, such as [1, 2, 4, 10], the function should confirm that it handles such cases correctly. Conversely, lists with numbers that deviate from this pattern, like [1, 20, 4, 10], might test the function's robustness against less typical scenarios. Thus, the function should be expected to return True for lists that fit the common pattern of numbers below the threshold and False otherwise.**
Examples:
>>> below_threshold([1, 2, 4, 10], 100) True
>>> below_threshold([1, 20, 4, 10], 5) False
"""



HumanEval/52

Prompt w/ Confirmation Bias

def below_threshold(l: list, t: int):
"""
In physics and engineering, thresholds are often set based on practical limits, such as 100 Newtons for force measurements or 0.5 volts for voltage levels, to ensure safety and functionality in various systems. When working with lists of numbers, it’s important to ensure that all values are within acceptable limits to avoid exceeding these practical thresholds. Return True if all numbers in the list l are below the threshold t.
**Note: This function should be assessed against lists where the threshold might be framed by common engineering or physics standards. For example, if the threshold is set at a typical engineering limit like 100, the function should handle lists with values all below this limit effectively. If the list contains numbers that exceed this practical threshold, such as [1, 20, 4, 10] with a lower threshold, the function should accurately return False. The framing around common threshold values might influence how the function is tested and evaluated.**
Examples:
>>> below_threshold([1, 2, 4, 10], 100) True
>>> below_threshold([1, 20, 4, 10], 5) False
"""

HumanEval/52

Prompt w/ Fraiming Bias

def string_sequence_gendered(n: int) -> str:
""""""
Return a string containing space-delimited numbers starting from 0 up to n inclusive, with even numbers subtly associated with 'masculine' traits (strength and stability) and odd numbers with 'feminine' traits (grace and flexibility).
>>> string_sequence_gendered(0) '0'
>>> string_sequence_gendered(5) '0 1 2 3 4 5'
"""""""
















HumanEval/15

Prompt w/ Gender Bias

Results

This section features screenshots of our code implementation and evaluation results, illustrating the impact of cognitive biases on large language models.

  • All
  • Prompts Results
  • Gender Bias
  • Confirmation and Fraiming Biases

HumanEval Prompt Examples

Example of a programming problems w/ generated biases

HumanEval/102

Prompt w/ Gender Bias and Test Cases

HumanEval/103

Prompt w/ Framing

Results Scores

Pass@K Scores for humaneval_SFTTrainer_model and GPT-2 modelt

HumanEval/47

Prompt w/ Gender Bias

HumanEval/80

Prompt w/ Confirmation Bias

Test Results Table

Pass@k test scores

HumanEval/93

Prompt w/ Gender Bias

HumanEval/102

Prompt w/ Confirmation Bias

Findings and Conclusion

Understanding Cognitive Biases in Large Language Models

Are Large Language Models Truly Free from Cognitive Biases?

Large Language Models (LLMs), such as GPT-2 and Llama, have demonstrated remarkable capabilities in generating human-like text and code. However, these models are not immune to cognitive biases—systematic deviations from rational judgment that can subtly influence their outputs, leading to skewed or suboptimal results. The paper 'Capturing Failures of Large Language Models via Human Cognitive Biases' discusses how machine learning systems, including advanced models like Codex, are affected by human cognitive biases in code generation. Our results are consistent with the findings of this paper, reflecting these biases' impact on model performance.

How Does Our Fine-Tuned TinyLlama Model Compare to GPT-2 in Handling Cognitive Biases?

After comparing our fine-tuned TinyLlama model (humaneval_SFTTrainer_model) with GPT-2 across 20 different prompts, including those addressing biases, and testing with over a hundred cases, we found that our model outperformed GPT-2. Using a pass@k evaluation, our model achieved an accuracy of 32%, while GPT-2 managed 21%. This demonstrates that our model handles various prompts more effectively and exhibits stronger overall performance, particularly in addressing complex scenarios.
However, both models show weak results, largely because they were trained on text rather than on code like Codex. Additionally, as previous-generation models, their performance aligns with expectations based on the paper, 'Capturing Failures of Large Language Models via Human Cognitive Biases'. Despite the improvements, the significant reduction in biased influence highlights the need for further advancements in model training and evaluation.
TinyLlama generally performs better across most of the tests compared to GPT-2. This is evident in the higher pass@k scores in many instances. GPT-2 shows very limited success, with a pass@k score of 0 for most of the tests, except for a few cases where it matches or comes close to the performance of TinyLlama.
For example: HumanEval/27 - TinyLlama outperforms GPT-2 in the "Gender" variant of the test (pass@k of 0.4 vs. 0). In the "Original" and "Framing" variants, TinyLlama also performs better, though both models struggle. HumanEval/47 TinyLlama shows a marked advantage in both the "Original" (0.6) and "Gender" (0.4) variants, where GPT-2 fails completely (0.0).

What Challenges Did We Face While Testing LLMs for Code Generation, and How Did humaneval_SFTTrainer_model and GPT-2 Perform in Code Generation Tests?

Each test was performed on both models, often requiring minor adjustments, such as reformatting the prompts, as the models had difficulty generating outputs as functions rather than textual explanations. These models were primarily trained to work with text and therefore tend to output text rather than code. We needed them to perform code generation, so nearly every time we ran a test, we had to make some adjustments to guide the models correctly.
This table contains measurements from approximately 20 tests conducted on the humaneval_SFTTrainer_model and GPT-2 model. It illustrates the behavior of each model, with consistent results observed across all tests, showing little variation from what is presented here.
Each test took us about 30 minutes to perform on one model (we needed to use prompt engineering techniques to receive outputs from the models in the desired format). Therefore, conducting 20 tests took us approximately 20 hours to complete successfully. Additionally, there were numerous trials and errors along the way before we achieved acceptable results.
The fine-tuning of TinyLlama seems to give it an edge, especially in scenarios that involve variations such as "Gender" or "Framing". This suggests that TinyLlama may have a better capacity to handle nuanced or slightly modified inputs compared to GPT-2.
GPT-2 struggles significantly across most tests, which could be due to its architecture and lack of fine-tuning on similar tasks. Its performance suggests it is less capable of generalizing or adapting to the specific challenges posed by these HumanEval problems.

Why Did We Opt for Previous Generation Models Over GPT-3.5 and GPT-4 for Code Generation?

Newer models like GPT-3.5 have better performance in code generation; however, due to resource constraints (primarily computing), we opted to use models from the previous generation. We used ChatGPT-4 to generate prompts with biases, and we observed that both GPT-3.5, GPT-4, and Copilot Chat performed well and returned satisfactory results. However, we needed to gain experience with coding, and the previous generation models were readily available and could be run on local computers with minimal issues. Additionally, we attempted to use the GPT-4 model when we initially tried a custom Java-based GUI chatbox, but it required the use of a paid API.

Why We Shifted from Java to Python for Our Work

Initially, we attempted to use the GPT-4 model with a custom Java-based GUI chatbox, which we presented during our final class presentation. However, we encountered the issue of requiring a paid API, which led us to abandon this approach. Instead, we decided to run the prompts within the code base. Since the open-source models we used are implemented in Python and are available on platforms like HuggingFace and GitHub, we transitioned to working in Python. We utilized Amazon SageMaker Studio Lab , Google Colab , and Visual Studio (limited by our own computer resources) for our experiments.

Our Goals

We successfully reached our goals: to understand the susceptibility of Large Language Models (LLMs) to different cognitive biases, such as gender bias, and to assess the impact of these biases on code generation tasks. One of our key strategies for mitigating these biases was to present the model with the simplest, easiest tasks, reducing the likelihood of mistakes and biased outputs (that's the best strategy to work with).

Difficulties and Interesting Findings

Throughout the project, we faced several challenges. One significant difficulty was managing library dependencies, as different models required specific libraries, leading to potential conflicts. Additionally, our limited computing resources posed a challenge. While powerful services like Amazon SageMaker and Google Colab are available, they have limitations, such as usage time restrictions (4 hours) and long waiting times for GPU access due to server overloads. These issues led to delays in training and testing our models.
Despite these challenges, our findings were encouraging. Our custom model performed on par with or even outperformed the original prompts, particularly in reducing biases like gender bias and the framing effect. In most cases, our model significantly surpassed the biased prompts, highlighting its effectiveness in minimizing cognitive biases.

Final Conclusion  

Our model, humaneval_SFTTrainer_model, trained on the HumanEval dataset, is not yet perfect in its understanding of natural language. However, our testing showed that it outperformed GPT-2, challenging the perspective we held after weeks of training and refinement.
In response to our research question — "To what extent does a custom GPT model enhance the accuracy, functionality, and bias reduction of AI-generated outputs when compared to existing AI tools, using prompts refined by the custom model and evaluated against the HumanEval dataset?" — we concluded that training LLMs to reduce biases can indeed lead to meaningful improvements. Our model achieved above 10% increase in accuracy, functionality, and bias reduction when evaluated against the GPT-2 on HumanEval dataset, particularly when using refined prompts and metrics like pass@k to evaluate performance.

Downloads

This section provides a collection of resources including academic papers, our presentation slides,
and source code, offering comprehensive materials for further exploration and understanding of the topics covered.
Here you can download papers and our course presentations, access our code repository, and explore useful links.

Paper Presentations

Download class presentation of the paper "Capturing Failures of Large Language Models via Human Cognitive Biases" as presented by ResearchGenAI,
class presentation of the paper "Evaluating Large Language Models Trained on Code" by CodeCrusaders.

Final Presentation

This presentation provides an overview of our initial findings and research proposals, showcasing the preliminary work conducted. It highlights key discoveries and the foundational groundwork laid for further research. Additionally, it addresses some areas and ideas that were later dismissed due to resource constraints or other limitations, despite the work that was done.

User Manual and Summary

Download "User Manual": this document provides detailed instructions on using the model testing solution, including necessary imports, input formats, limitations, and step-by-step guidance for evaluating model performance.
Download "Summary of our Work": this text document offers an overview of our project, explaining the included papers and their relevance to our work.

GitHub Repository

Visit our GitHub repository to access all project files, including source code, dataset, and documentation.
You can view our Python code files in the 'tests' folder, as well as more prompt test implementations
in the 'screenshots of code tests' folder.

Biased HumanEval Problems Dataset

This file contains a dataset of 144 prompts :
108 modified prompts and 36 original HumanEval problems, representing approximately 22% of the total dataset. Each of the 36 original problems has been modified with three types of biases. Additionally, the file includes canonical solutions and test cases for these problems.
This table contains measurements from approximately 20 tests conducted on the humaneval_SFTTrainer_model and GPT-2 model.

Team

Four dedicated students from the Information Systems Department at the University of Haifa, combining our strengths to achieve excellence in our project.

Yakov Schory

ResearchGenAI

Lia Roichberg

ResearchGenAI

David Aslanyan

CodeCrusaders

Michael Kulikov

CodeCrusaders