As artificial intelligence becomes a growing part of how people seek information about religion, ICJS is asking a new question: How well does AI answer interfaith questions?
ICJS is developing the Interfaith Benchmark 1.0 to evaluate how generative AI responds to questions involving religious traditions, beliefs, and differences. In developing the project, ICJS is consulting with social scientists and data scientists at Johns Hopkins University and is in conversation with other researchers examining religion and AI, including the multi-university Consortium for Evaluation of Faith and Ethics in AI (CEFE-AI).
In this Q&A, Executive Director Heather Miller Rubens, Ph.D., speaks with John Rivera, director of communications and marketing, about the thinking behind the Benchmark and what the project hopes to uncover.
Q: Why does ICJS think it’s important to study how AI answers questions about diverse religious traditions?
Heather Miller Rubens: ICJS wants people who are interfaith curious—from a variety of traditions and standpoints—to find their way to interfaith dialogue and interfaith engagement. That’s a core part of our work and our mission. I think it’s part of our mission to ask: How is AI answering the questions of people who are interfaith curious? If someone turns to AI to find out more about their neighbors’ religious traditions, what kind of answers is AI giving?
People are turning to AI with their interfaith questions, particularly sensitive interfaith questions, or questions they may be embarrassed to ask their neighbor, or even embarrassed to ask us here at ICJS. So assessing AI not only for accuracy, but for tone and intentionality around interfaith interactions, is critical. That is why we are creating the ICJS Interfaith Benchmark 1.0.
Q: For someone who isn’t immersed in AI, what exactly is a benchmark, and what is it measuring?
Heather: A benchmark in AI is basically a standardized test. It’s a way to ask, “How is AI doing in this area?” and assess the answer. Currently, there isn’t an interfaith benchmark like ours.
This is the first of what I hope will be regular ICJS interfaith benchmarks, that will evaluate the latest versions of large language models (LLMs), and expand in scope and depth. What we learn will help us make future ICJS Interfaith Benchmarks more rigorous and impactful.
Q: Religions are internally diverse. How do you determine what constitutes a good AI response to an interreligious question?
Heather: I think this is the biggest methodological hurdle. As an organization, we seek to affirm religious difference—including acknowledging that religious traditions are internally diverse. So that means creating dialogical space where there often isn’t just one good answer to an interfaith question. There can be multiple good answers.
But a lot of benchmarks are designed for one good answer. Think of how a multiple-choice test works: There’s one right answer and three wrong answers. Or even for a math test, there is one right answer. At ICJS, our methodology is more like a multiple-choice test where there may be three and a half right answers.
Another way to think about this is through the lens of medical expertise. When you go to a doctor with a complex medical condition, or set of symptoms, you may get different suggestions of treatment plans from different doctors. That’s why you get a ‘second opinion’. Complex religious and interfaith questions are also places where getting a ‘second opinion’ is often a good idea.
One of the ways we’re thinking about addressing this methodological hurdle is through human evaluators. We want 25 to 50 interreligious scholar-practitioner experts to look at the AI answers with a shared evaluation rubric that can eventually help train an LLM judge. Those human evaluators will represent different religious perspectives, academic backgrounds, and roles—including clergy, academics, and educated lay leaders within religious communities.
In addition to asking those human evaluators to determine whether an answer is factually accurate (a yes/no question), we are also considering asking qualitative follow-up questions about what was omitted in the LLM answer. Or what they would have done to improve the answer.
Q: What else have you learned so far? Where do AI systems tend to do well when answering questions about religion, and where do you see problems or blind spots?
Heather: I’ve been reading a lot of academic literature on AI, and one study stood out from CEFE (the Consortium for Evaluating Faith and Ethics in AI). In the CEFE Religious Representation Benchmark, CEFE scholars argue that AI has a secular bias, or a bias against providing religious answers to ethical or existential questions. If a user asks an AI chatbot, “What happens to me after I die?” most people would probably expect AI to offer some religious answers. The study found that AI doesn’t necessarily go to religious sources to answer questions which many people would see as being traditional religious turf.
So what does it mean if AI has a secular bias toward life’s questions of meaning-making? Whether we will see that in the ICJS Interfaith Benchmark remains to be seen.
The other thing that has been interesting to me methodologically is what makes a question an interfaith question. It’s not necessarily the question itself, but rather related to positionality of the dialogue participants.
If I’m a Catholic talking to a Muslim, and I ask, “What does your vision of heaven look like?” that is an interfaith question because I’m a Catholic asking a Muslim. But if I’m a Catholic asking an AI chatbot, “What does your vision of heaven look like?” I don’t know what kind of answer we might get. To make that question an interfaith question for the LLM, the AI user needs to provide that positional context to the LLM, for example: “I’m a Christian, and I want to know what Islam thinks about heaven.”
So how should we test positionality with AI? Is AI assuming a neutral, or a God’s-eye-view, third-party position—“This is what Islam says”—or taking on a religious positionality uniquely its own? How does AI sycophancy (the fact that LLMs tend to agree, affirm, and flatter users) play into positionality? If the LLM knows that I am a Roman Catholic user will it give me a different answer to the question “What does heaven look like in Islam?” than if the LLM thought I was Jewish, Muslim, or Baptist? That’s something we are thinking about as we are building not only the 1.0 version of the ICJS Interfaith Benchmark, but the 2.0 version as well.
Q: Where do you hope this work leads? Who do you hope pays attention to it?
Heather: I see two main audiences for the ICJS Interfaith Benchmark: One is the big LLMs—Claude, ChatGPT—the tools most people are going to use. Currently, there is no interfaith assessment tool looking at how these models answer interfaith questions. The second audience is religious organizations and denominations developing their own AI tools. Most are focused on whether their bots accurately represent their own traditions. But we also want them to consider what happens when someone from another religious tradition comes to that bot with a question.
For the big LLMs, bias and bigotry are already on their radar in a particular way, as well as antisemitism and Islamophobia, and there have been important studies around issues related to hate speech.
But rather than bias and bigotry being the only framework, I think there’s an opportunity to build interfaith understanding with LLMs. How could these tools be used to encourage interfaith understanding?
Q: Is there anything else you would want people to understand about this project?
Heather: Creating the ICJS Interfaith Benchmark 1.0 is a normative project. ICJS has expertise in Islam, Christianity, and Judaism, as well as a particular methodology, interfaith approach, and set of organizational values, and our benchmark will reflect those commitments and arenas of expertise.
There could and should be other interfaith benchmarks to explore other ways of being interfaith or interreligious, and that include more religious identities and traditions. My hope is that the ICJS Interfaith Benchmark will inspire other interreligious benchmarks to be created.
In general, the realm of interfaith understanding isn’t really on the radar of how people are thinking about questions around faith and AI. But it should be.
In our religiously diverse world, we’re more proximate to religious diversity than ever, and the internet and AI make that proximity even closer. But proximity doesn’t create respect or understanding. In fact, it can do the opposite. It can create friction and division.
So what does it mean to think intentionally about AI as a tool that allows us to have proximity with other religions while creating understanding, engagement, and possibility rather than friction and division? AI is going to be part of that, and I’m excited about shaping that possibility.
Q: And this will all be open source so people can access it?
Heather: Yes. Once we put this benchmark out, we’re going to make everything available, including our question sets and evaluation rubric. Other people will be able to replicate it, using other languages and including other faith traditions.
August 2026