CodeGenCrusaders
Unveiling the Impact of Cognitive Biases on Large Language Models
Cognitive Biases
Cognitive bias is a systematic pattern of deviation from norm or rationality in judgment. Individuals create their own "subjective reality" from their perception of the input.
Framing Bias
Framing effect or bias is a cognitive bias in which people decide between options based on whether the options are presented with positive or negative connotations.
Confirmation Bias
Confirmation bias is the tendency to search for, interpret, favor, and recall information in a way that confirms or supports one's prior beliefs or values.
Gender Bias
Gender bias is a widespread set of implicit biases that discriminate against a gender or prefer one gender over the other, which leads to treating individuals unequally and unfairly based on their gender.
About Our Project
Welcome to CodeGenCrusaders, a project conducted within the framework of a
seminar course titled "Software Engineering in the Age of AI" at Haifa University.
This website shares our findings and insights.
Project Overview
Large Language Models (LLMs) like GPT-2 and Llama have become powerful tools in software development, capable of generating code, completing tasks, and assisting developers in various ways. However, these models are not immune to the pitfalls of cognitive biases, which can lead to skewed or suboptimal outputs. Our project seeks to explore how LLMs are influenced by biased prompts, particularly focusing on framing, confirmation, and gender biases.
Research Question
To what extent does a custom GPT model enhance the accuracy, functionality, and bias reduction of AI-generated outputs when compared with existing AI tools, using prompts refined by the custom model and evaluated against the HumanEval dataset?
Research Foundation
Our work is based on two seminar papers:
- "Capturing Failures of Large Language Models via Human Cognitive Biases"
- "Evaluating Large Language Models Trained on Code"
Drawing from these studies, we designed a series of experiments using the HumanEval dataset as our benchmark. We created a variety of biased prompts that introduced different types of cognitive biases with help of ChatGPT-3.5 and 4, and then tested these prompts on different LLMs, mainly GPT-2 model.
The TinyLlama project aimed to pretrain a 1.1B Llama model on 3 trillion tokens, while GPT-2 was pretrained on 8 million web pages. In our project, we used these models to examine how cognitive biases in prompts affect their responses, revealing the impact of biased input on output accuracy and reliability.
Work Process and Methodology
Methodology
- Bias Types: framing bias, confirmation bias, and gender bias.
- Models Tested: GPT-2 and TinyLlama models.
- Benchmark Dataset: HumanEval.
- Evaluation Metric: pass@k metric.
Our experiments were conducted locally on our computers to ensure controlled conditions and reproducible results. The outcomes offer valuable insights into how various biases can impact the performance and reliability of LLM-generated code.
Project Goals
- To understand the susceptibility of LLMs to different cognitive biases.
- To assess the impact of these biases on code generation tasks.
- To propose potential mitigation strategies for developers and AI practitioners.
Prompts Containing Cognitive Biases
Dataset and Bias Implementation
We utilized HumanEval problems as our dataset and benchmark. From this dataset, we selected 36 problems, representing approximately 22% of the total dataset. We then created a total of 108 modified problems with biases (three types of biases applied to each of the 36 original problems) and retained the original 36 problems for comparison. Using ChatGPT-4, we introduced framing, confirmation, and gender biases into these problems.
Resource Preparation and Refinement
Prior to this, we provided ChatGPT with the paper "Capturing Failures of Large Language Models via Human Cognitive Biases" and examples of various biases, as well as the paper "Uncovering and Quantifying Social Biases in Code." We then refined the biased prompts as needed. Due to our limited background in educational psychology and sociology, we relied heavily on ChatGPT, which demonstrated strong performance in cognitive techniques during our seminar.
Examples of HumanEval problems modified to highlight cognitive biases and their impact on coding tasks. See how these biases influence problem interpretation.
Results
This section features screenshots of our code implementation and evaluation results, illustrating the impact of cognitive biases on large language models.
- All
- Prompts Results
- Gender Bias
- Confirmation and Fraiming Biases
Findings and Conclusion
Understanding Cognitive Biases in Large Language Models
Are Large Language Models Truly Free from Cognitive Biases?
Large Language Models (LLMs), such as GPT-2 and Llama, have demonstrated remarkable capabilities in generating human-like text and code. However, these models are not immune to cognitive biases—systematic deviations from rational judgment that can subtly influence their outputs, leading to skewed or suboptimal results. The paper 'Capturing Failures of Large Language Models via Human Cognitive Biases' discusses how machine learning systems, including advanced models like Codex, are affected by human cognitive biases in code generation. Our results are consistent with the findings of this paper, reflecting these biases' impact on model performance.
How Does Our Fine-Tuned TinyLlama Model Compare to GPT-2 in Handling Cognitive Biases?
After comparing our fine-tuned TinyLlama model (humaneval_SFTTrainer_model)
with GPT-2 across 20 different prompts, including those addressing biases,
and testing with over a hundred cases, we found that our model outperformed GPT-2.
Using a pass@k evaluation, our model achieved an accuracy of 32%, while GPT-2
managed 21%. This demonstrates that our model handles various prompts more
effectively and exhibits stronger overall performance, particularly in
addressing complex scenarios.
However, both models show weak results, largely because they were trained on
text rather than on code like Codex. Additionally, as previous-generation models,
their performance aligns with expectations based on the paper, 'Capturing Failures
of Large Language Models via Human Cognitive Biases'. Despite the improvements,
the significant reduction in biased influence highlights the need for further
advancements in model training and evaluation.
TinyLlama generally performs better across most of the tests compared to GPT-2.
This is evident in the higher pass@k scores in many instances. GPT-2 shows very
limited success, with a pass@k score of 0 for most of the tests, except for a
few cases where it matches or comes close to the performance of TinyLlama.
For example: HumanEval/27 - TinyLlama outperforms GPT-2 in the "Gender" variant
of the test (pass@k of 0.4 vs. 0). In the "Original" and "Framing" variants,
TinyLlama also performs better, though both models struggle. HumanEval/47 TinyLlama
shows a marked advantage in both the "Original" (0.6) and "Gender" (0.4) variants,
where GPT-2 fails completely (0.0).
What Challenges Did We Face While Testing LLMs for Code Generation, and How Did humaneval_SFTTrainer_model and GPT-2 Perform in Code Generation Tests?
Each test was performed on both models, often requiring minor adjustments,
such as reformatting the prompts, as the models had difficulty
generating outputs as functions rather than textual explanations.
These models were primarily trained to work with text and therefore
tend to output text rather than code. We needed them to perform code
generation, so nearly every time we ran a test, we had to make some
adjustments to guide the models correctly.
This table contains measurements from approximately 20 tests
conducted on the humaneval_SFTTrainer_model and GPT-2 model.
It illustrates the behavior of each model, with consistent results
observed across all tests, showing little variation from what is
presented here.
Each test took us about 30 minutes to perform on one model
(we needed to use prompt engineering techniques to receive outputs
from the models in the desired format). Therefore, conducting 20
tests took us approximately 20 hours to complete successfully.
Additionally, there were numerous trials and errors along the way
before we achieved acceptable results.
The fine-tuning of TinyLlama seems to give it an edge, especially
in scenarios that involve variations such as "Gender" or "Framing".
This suggests that TinyLlama may have a better capacity to handle
nuanced or slightly modified inputs compared to GPT-2.
GPT-2 struggles significantly across most tests, which could be due
to its architecture and lack of fine-tuning on similar tasks.
Its performance suggests it is less capable of generalizing or
adapting to the specific challenges posed by these HumanEval problems.
Why Did We Opt for Previous Generation Models Over GPT-3.5 and GPT-4 for Code Generation?
Newer models like GPT-3.5 have better performance in code generation; however, due to resource constraints (primarily computing), we opted to use models from the previous generation. We used ChatGPT-4 to generate prompts with biases, and we observed that both GPT-3.5, GPT-4, and Copilot Chat performed well and returned satisfactory results. However, we needed to gain experience with coding, and the previous generation models were readily available and could be run on local computers with minimal issues. Additionally, we attempted to use the GPT-4 model when we initially tried a custom Java-based GUI chatbox, but it required the use of a paid API.
Why We Shifted from Java to Python for Our Work
Initially, we attempted to use the GPT-4 model with a custom Java-based GUI chatbox, which we presented during our final class presentation. However, we encountered the issue of requiring a paid API, which led us to abandon this approach. Instead, we decided to run the prompts within the code base. Since the open-source models we used are implemented in Python and are available on platforms like HuggingFace and GitHub, we transitioned to working in Python. We utilized Amazon SageMaker Studio Lab , Google Colab , and Visual Studio (limited by our own computer resources) for our experiments.
Our Goals
We successfully reached our goals: to understand the susceptibility of Large Language Models (LLMs) to different cognitive biases, such as gender bias, and to assess the impact of these biases on code generation tasks. One of our key strategies for mitigating these biases was to present the model with the simplest, easiest tasks, reducing the likelihood of mistakes and biased outputs (that's the best strategy to work with).
Difficulties and Interesting Findings
Throughout the project, we faced several challenges. One significant difficulty was managing library
dependencies, as different models required specific libraries, leading to potential conflicts.
Additionally, our limited computing resources posed a challenge. While powerful services like Amazon
SageMaker and Google Colab are available, they have limitations, such as usage time restrictions (4 hours)
and long waiting times for GPU access due to server overloads. These issues led to delays in training and
testing our models.
Despite these challenges, our findings were encouraging. Our custom model performed on par with or even
outperformed the original prompts, particularly in reducing biases like gender bias and the framing effect.
In most cases, our model significantly surpassed the biased prompts, highlighting its effectiveness
in minimizing cognitive biases.
Final Conclusion
Our model, humaneval_SFTTrainer_model, trained on the HumanEval dataset,
is not yet perfect in its understanding of natural language.
However, our testing showed that it outperformed GPT-2, challenging the perspective
we held after weeks of training and refinement.
In response to our research question — "To what extent does a custom GPT model enhance the accuracy,
functionality, and bias reduction of AI-generated outputs when compared to existing AI tools,
using prompts refined by the custom model and evaluated against the HumanEval dataset?" —
we concluded that training LLMs to reduce biases can indeed lead to meaningful improvements.
Our model achieved above 10% increase in accuracy, functionality, and bias reduction when evaluated
against the GPT-2 on HumanEval dataset, particularly when using refined prompts and metrics like pass@k
to evaluate performance.
Downloads
This section provides a collection of resources including academic
papers, our presentation slides,
and source code, offering comprehensive materials for further
exploration and understanding of the topics covered.
Here you can download papers and our course presentations, access our code repository,
and explore useful links.
Paper Presentations
Download class presentation of the paper
"Capturing Failures of Large Language Models via Human Cognitive Biases"
as presented by ResearchGenAI,
class presentation of the paper
"Evaluating Large Language Models Trained on Code"
by CodeCrusaders.
Final Presentation
This presentation provides an overview of our initial findings and research proposals, showcasing the preliminary work conducted. It highlights key discoveries and the foundational groundwork laid for further research. Additionally, it addresses some areas and ideas that were later dismissed due to resource constraints or other limitations, despite the work that was done.
Study Papers from arXiv That Shaped Our Research
"Capturing Failures of Large Language Models via Human Cognitive Biases",
"Evaluating Large Language Models Trained on Code",
"Uncovering and Quantifying Social Biases in Code Generation"
User Manual and Summary
Download
"User Manual":
this document provides detailed instructions on using the model testing solution,
including necessary imports, input formats, limitations, and step-by-step guidance
for evaluating model performance.
Download
"Summary of our Work": this text document offers an overview of our project,
explaining the included papers and their relevance to our work.
GitHub Repository
Visit our GitHub repository to access all project files,
including source code, dataset, and documentation.
You can view our Python code files in the 'tests' folder,
as well as more prompt test implementations
in the 'screenshots of code tests' folder.
Biased HumanEval Problems Dataset
This file contains a dataset of 144 prompts
:
108 modified prompts and 36 original HumanEval problems,
representing approximately 22% of the total dataset. Each of the 36 original problems
has been modified with three types of biases. Additionally, the file includes canonical
solutions and test cases for these problems.
This table contains measurements
from approximately 20 tests
conducted on the humaneval_SFTTrainer_model and GPT-2 model.
Team
Four dedicated students from the Information Systems Department at the University of Haifa, combining our strengths to achieve excellence in our project.
Yakov Schory
ResearchGenAI
Lia Roichberg
ResearchGenAI
David Aslanyan
CodeCrusaders