PolyInterview: An LLM-based Platform for Immersive Mock Interview Practice with Comprehensive Multimodal Assessment
Summary
This paper presents PolyInterview, an LLM-based platform that conducts immersive mock interviews with a lip-synced digital human, providing comprehensive multimodal assessment of response content, vocal delivery, and non-verbal behavior. The system generates role-tailored questions and evaluates candidates using 13 behavioral features linked to KSA and STAR frameworks.
View Cached Full Text
Cached at: 07/14/26, 04:21 AM
# An LLM-based Platform for Immersive Mock Interview Practice with Comprehensive Multimodal Assessment
Source: [https://arxiv.org/html/2607.10310](https://arxiv.org/html/2607.10310)
Zhiyuan Wen1, Jiannong Cao1, Zijian Wang1, Chen Chen1, Xiaoyun Liu1, Jianing Yin1, Zhuo Li2 1Department of Computing, The Hong Kong Polytechnic University, Hong Kong, China 2School of Artificial Intelligence, Chongqing University of Posts and Telecommunications, China \{zyuanwen, jiannong\.cao, zi\-jian\.wang\}@polyu\.edu\.hk \{chen03\.chen, xiaoyun\.liu, jianing\-laetitia\.yin\}@connect\.polyu\.hk cliff\.zhuo\.li@gmail\.com
###### Abstract
Preparing for job interviews is important for securing desired positions, yet realistic practice remains difficult to access: real interviews are infrequent, expert mock coaching is costly, and self\-practice offers neither adaptive dialogue nor structured assessment\. Existing systems typically address only parts of this need through fixed question sequences, limited communication channels, or feedback with little supporting evidence\. We presentPolyInterview, an LLM\-based platform for immersive mock interview practice with comprehensive multimodal assessment\. PolyInterview uses the target job description and CV to generate questions tailored to the role and candidate, conducts multi\-turn spoken interviews with a lip\-synced digital human interviewer that asks answer\-aware follow\-up questions, and evaluates response content, vocal delivery, and non\-verbal behavior\. Four parallel evaluators produce 13 behavior\-level features that are aggregated into 10 assessment aspects and two competency tracks\. Guided by the KSA and STAR frameworks, the report links each score to behavioral evidence and actionable recommendations\. PolyInterview is publicly accessible111PolyInterview is deployed for public access\. The demo video and system link are available[here](https://dannywang1922.github.io/polyinterview)\.\. Its current all\-account snapshot contains 101 accounts, 1,564 interview sessions, 7,665 generated questions, and 1,422 five\-stage question sets\. Generated questions are more closely aligned with their matched job description than with cross\-role job descriptions in 93\.7% of sessions\. An evaluation by ten experts found strong question plans and actionable feedback\.
PolyInterview: An LLM\-based Platform for Immersive Mock Interview Practice with Comprehensive Multimodal Assessment
Zhiyuan Wen1, Jiannong Cao1, Zijian Wang1, Chen Chen1, Xiaoyun Liu1, Jianing Yin1, Zhuo Li21Department of Computing, The Hong Kong Polytechnic University, Hong Kong, China2School of Artificial Intelligence, Chongqing University of Posts and Telecommunications, China\{zyuanwen, jiannong\.cao, zi\-jian\.wang\}@polyu\.edu\.hk\{chen03\.chen, xiaoyun\.liu, jianing\-laetitia\.yin\}@connect\.polyu\.hkcliff\.zhuo\.li@gmail\.com
Figure 1:PolyInterview’s workflow, from personalized setup and immersive interviewing to multimodal assessment and comprehensive reporting and feedback\.## 1Introduction
Effective preparation for job interviews is critical to securing a desired position\. During an interview, both candidates’ role\-relevant knowledge and skills and their ability to communicate experience and reasoning clearly are evaluated\. Despite the importance of interview preparation, realistic practice remains difficult to access: real interviews are infrequent, expert mock coaching is costly, and self\-practice offers neither adaptive dialogue nor structured assessment\. These constraints motivate an accessible platform that can simulate multi\-turn interview interactions and provide comprehensive assessment and feedback\.
Commercial products and research prototypes have expanded access to interview practice, yet, as summarized in Table[1](https://arxiv.org/html/2607.10310#S1.T1), they do not jointly support role\-conditioned multi\-stage questioning, answer\-aware probing, and theory\-grounded assessment of textual, vocal, and non\-verbal behavior within an integrated practice environment\. For commercial tools, products such as Google Interview Warmup, Yoodli, and Big Interview support question rehearsal or automated feedback, but generally rely on fixed question sequences, assess only selected communication channels, or provide limited transparency into how feedback is derived\. For research prototypes, virtual interview agents have explored conversational realism\(Hoque et al\.,[2013](https://arxiv.org/html/2607.10310#bib.bib7); Anderson et al\.,[2013](https://arxiv.org/html/2607.10310#bib.bib1); Smith et al\.,[2014](https://arxiv.org/html/2607.10310#bib.bib17)\); LLM\-based mock interviewers support flexible question generation\(Li et al\.,[2023](https://arxiv.org/html/2607.10310#bib.bib11); Sun et al\.,[2025](https://arxiv.org/html/2607.10310#bib.bib18); Daryanto et al\.,[2025](https://arxiv.org/html/2607.10310#bib.bib4)\); and automated video interview systems analyze non\-verbal behavior\(Naim et al\.,[2018](https://arxiv.org/html/2607.10310#bib.bib13); Hemamou et al\.,[2019](https://arxiv.org/html/2607.10310#bib.bib5)\)\.
To address these limitations, we presentPolyInterview, an LLM\-based platform for immersive mock interview practice with comprehensive multimodal assessment \(Figure[1](https://arxiv.org/html/2607.10310#S0.F1)\)\. PolyInterview personalizes each mock interview by using the target job description \(JD\) and the candidate’s CV to generate tailored questions\. A lip\-synced digital human interviewer creates an immersive interview setting, while LLM\-based agents sustain multi\-turn interaction and generate answer\-aware follow\-up questions\. For comprehensive assessment, we combine LLM\-based analysis of response content, speech analysis, and vision\-language model \(VLM\)\-based analysis of non\-verbal behavior\. Guided by the Knowledge, Skills, and Abilities \(KSA\) framework\(Peterson et al\.,[2001](https://arxiv.org/html/2607.10310#bib.bib15)\)and the Situation–Task–Action–Result \(STAR\) structure for organizing evidence in patterned behavior description interviews\(Janz,[1982](https://arxiv.org/html/2607.10310#bib.bib9)\), the assessment captures two complementary dimensions of interview performance: the extent to which candidates demonstrate knowledge, skills, and experience relevant to the target role, and how clearly and effectively they communicate them\. The resulting 13 behavior\-level features are aggregated into 10 assessment aspects and two competency tracks, with each score linked to behavioral evidence and actionable feedback\.
PolyInterview’s current all\-account snapshot contains 101 accounts and 1,564 interview sessions\. It includes 7,665 generated questions spanning 83 position titles, with 1,422 session question sets covering the complete five\-stage interview structure\. A lexical alignment analysis further shows that generated question sets are more closely aligned with their matched JD than with cross\-role JDs in 93\.7% of sessions, and the matched JD ranks first in 82\.4% of sessions\. An evaluation by ten human experts further identified strong question plan quality and actionable recommendations, while revealing follow\-up diagnostic depth and response faithfulness as areas for improvement\. In summary, our contributions are threefold:
1. 1\.Immersive, personalized interview simulation\.We develop an end\-to\-end workflow that constructs a multi\-stage interview from the JD and CV, adapts follow\-up questions to candidate responses, and delivers the interaction through a configurable, lip\-synced digital human interviewer \(§[2\.1](https://arxiv.org/html/2607.10310#S2.SS1)–§[2\.2](https://arxiv.org/html/2607.10310#S2.SS2)\)\.
2. 2\.Comprehensive multimodal assessment\.We introduce a pipeline grounded in KSA and STAR that analyzes textual, vocal, and visual behavior through four evaluators, maps 13 features to 10 aspects and two competency tracks, and preserves a traceable path from each score to actionable feedback \(§[2\.3](https://arxiv.org/html/2607.10310#S2.SS3)\)\.
3. 3\.Deployment and expert evaluation\.We characterize an all\-account snapshot of 101 accounts and 1,564 sessions, analyze workflow coverage and role alignment, and evaluate question plans, follow\-ups, and feedback reports with ten human experts \(§[4](https://arxiv.org/html/2607.10310#S4)and §[5](https://arxiv.org/html/2607.10310#S5)\)\.
Table 1:Comparison with representative commercial and research interview systems \(✓full,∼\\simpartial,×\\timesabsent\)\.
## 2System Overview and Design
Figure 2:PolyInterview’s pipeline for personalized questioning, answer\-aware follow\-ups, multimodal assessment, and comprehensive feedback\.A PolyInterview session comprises four stages \(Figure[2](https://arxiv.org/html/2607.10310#S2.F2)\):*personalized setup*,*immersive interview*,*multimodal assessment*, and*comprehensive report and feedback*\. The candidate configures a personalized session, interacts with the digital human interviewer through multi\-turn spoken dialogue, and reviews both per\-question and whole\-interview feedback\. We first describe the interface and user workflow \(§[2\.1](https://arxiv.org/html/2607.10310#S2.SS1)\), followed by the mechanisms for adaptive interviewing \(§[2\.2](https://arxiv.org/html/2607.10310#S2.SS2)\) and comprehensive multimodal assessment \(§[2\.3](https://arxiv.org/html/2607.10310#S2.SS3)\)\.
### 2\.1User Interface and Workflow
\(a\)Personalized setup\.
\(b\)Immersive digital human interview\.
\(c\)Overall assessment summary\.
\(d\)Behavior\-level feature profile\.
\(e\)KSA\-aligned assessment aspect profile\.


\(f\)Per\-question diagnosis and improvement suggestions\.
Figure 3:PolyInterview’s interface for personalized setup, immersive interviewing, and multilevel assessment results\.#### Interview setup\.
The setup interface consolidates the inputs required for personalization on one screen \(Figure[3\(a\)](https://arxiv.org/html/2607.10310#S2.F3.sf1)\)\. The candidate selects one of three interviewer personas, specifies a target company, position, and JD, uploads a CV in PDF format, and chooses a session of 8, 15, or 30 minutes\. These inputs condition both the content and length of the interview\. The JD and CV provide the evidence used to tailor the interview to the role and candidate, while the selected duration determines the question budget\.
#### Immersive live interview\.
A lip\-synced digital human interviewer presents each question aloud, and streaming speech recognition supports spoken responses\. The interface places the interviewer and candidate video side by side and displays connection status, session progress, and response controls \(Figure[3\(b\)](https://arxiv.org/html/2607.10310#S2.F3.sf2)\)\. For each question, the system records the response transcript, audio, and video, enabling subsequent assessment of response content, vocal delivery, and non\-verbal behavior\.
#### Comprehensive assessment results\.
After the interview, the interface presents assessment results and feedback from summary to detail\. The overall report presents the overall assessment, two competency\-track scores, strengths, and improvement priorities \(Figure[3\(c\)](https://arxiv.org/html/2607.10310#S2.F3.sf3)\)\. Candidates can then inspect the 10 KSA\-aligned assessment aspects \(Figure[3\(e\)](https://arxiv.org/html/2607.10310#S2.F3.sf5)\) and drill down to the 13 behavior\-level features produced by the four evaluators \(Figure[3\(d\)](https://arxiv.org/html/2607.10310#S2.F3.sf4)\)\. Finally, the per\-question view pairs each prompt and response with an answer\-specific diagnosis and improvement suggestions, including STAR\-guided phrasing that turns the candidate’s experience into a structured account \(Figure[3\(f\)](https://arxiv.org/html/2607.10310#S2.F3.sf6)\)\. This hierarchy connects the assessment framework in §[2\.3](https://arxiv.org/html/2607.10310#S2.SS3)to concrete practice guidance\. The complete report can also be exported as a PDF\.
### 2\.2Adaptive Interviewing
Personalization operates both before and during the interview: the Interview Planner Agent combines role requirements from the target JD with relevant skills, projects, and experience from the candidate’s CV to generate questions tailored to the role and candidate, while the Interviewer Agent produces bounded, answer\-aware follow\-ups\. Appendix[C](https://arxiv.org/html/2607.10310#A3)provides the detailed question\-planning and follow\-up workflows\.
### 2\.3Comprehensive Multimodal Assessment
Figure 4:Three\-layer multimodal assessment from 13 behavior\-level features to 10 aspects and two competency tracks\.Figure[4](https://arxiv.org/html/2607.10310#S2.F4)presents the three\-layer assessment architecture\. Four evaluators analyze each response in parallel, and their outputs are progressively aggregated from behavior\-level features to assessment aspects and competency tracks\.
#### Layer 1: Behavior\-level features\.
Two LLM\-based text evaluators analyze response content and expression\. The*Professional Performance*evaluator assesses conceptual accuracy, terminology use, and problem\-solving logic, while the*Way of Expression*evaluator assesses structural clarity, coherence, and word usage\. A VLM\-based*Non\-verbal Behavior*evaluator analyzes eye contact, facial expression, body posture, and gesture from the recorded video\. A speech\-based*Oral Expression*evaluator analyzes pronunciation accuracy, prosody, and fluency\. Together, these evaluators produce 13 feature scores on a scale from 0 to 10\.
#### Layer 2: Assessment aspects\.
An aggregation agent maps the 13 features to 10 KSA\-aligned aspects in three families:*cognitive*\(general intelligence, applied mental skills, and creativity\),*background*\(job knowledge and skills, education and training, and experience and work history\), and*social*\(communication, interpersonal skills, leadership, and persuasiveness\)\. The social family is informed by the Big Five\(McCrae and Costa,[1987](https://arxiv.org/html/2607.10310#bib.bib12)\), without inferring or reporting personality traits\. Each aspect combines primary and secondary features with a 70/30 weighting\. Question\-category gating restricts assessment to aspects that the current question can reasonably elicit\. For example, behavioral questions activate interpersonal skills, applied mental skills, and leadership, whereas skill\-QA questions activate job knowledge, general intelligence, and education\. Appendix[E](https://arxiv.org/html/2607.10310#A5)details the mapping and gating rules\.
#### Layer 3: Competency tracks\.
The cognitive and background aspects are aggregated into the*Professional Competency*track, and the social aspects form the*Communication Competency*track\. The overall score is the mean of these two tracks\. A final CV\-informed step adapts the recommendations to the candidate’s experience and target role\.
#### Theory grounding and traceability\.
The KSA framework specifies the job\-relevant competencies represented by the professional track\(Peterson et al\.,[2001](https://arxiv.org/html/2607.10310#bib.bib15)\), while STAR structures the evidence expected in behavioral responses\(Janz,[1982](https://arxiv.org/html/2607.10310#bib.bib9)\)\. For example, an omitted Result can be converted into a targeted recommendation to state the outcome and its significance\. Each competency score can be traced through its contributing aspects to the behavior\-level features and modalities that produced it, linking theoretical criteria to inspectable evidence\. Appendix[D](https://arxiv.org/html/2607.10310#A4)details how KSA, STAR, and the Big Five enter the assessment pipeline\.
## 3Implementation and Deployment
#### Real\-time pipeline\.
The digital human interviewer is driven by lip\-sync generation\(Prajwal et al\.,[2020](https://arxiv.org/html/2607.10310#bib.bib16)\)and streamed to the browser over WebRTC, with streaming speech recognition and synthesis supporting the spoken interaction\. LLM\-based agents manage question generation and follow\-up reasoning, a VLM analyzes non\-verbal behavior\(Bai et al\.,[2023](https://arxiv.org/html/2607.10310#bib.bib2)\), and a speech service provides pronunciation analysis\. The platform runs as four cooperating services behind HTTPS\. A session pool caps concurrent digital human sessions at five and queues additional requests\. The four evaluators run concurrently for each response through a thread pool\. Model versions, service layout, and stack details are provided in Appendix[B](https://arxiv.org/html/2607.10310#A2)\.
#### Demo interaction\.
A visitor selects an interviewer persona and either uploads a CV or uses a bundled sample CV and JD\. After a brief spoken interview with three or four questions, the visitor can inspect the report from the two competency tracks through the assessment aspects to individual behavior\-level features and recommendations\. A prerecorded end\-to\-end session and exported PDF report provide a network\-independent demonstration fallback\.
## 4System Usage Analysis
Table 2:All\-account usage snapshot, including internal and test activity\.#### Usage scale and workflow coverage\.
PolyInterview runs on the university network\. The snapshot includes 101 account directories and 1,564 sessions, including test, internal, placeholder, and team accounts, so these totals characterize platform activity rather than verified external candidates\. It contains 7,665 generated questions across 83 position titles, and 1,422 session question sets \(90\.9%\) cover all five planned categories: self\-introduction, behavioral, skill\-QA, scenario, and candidate questions \(Table[2](https://arxiv.org/html/2607.10310#S4.T2)\)\. The snapshot also contains 1,744 WAV and 1,744 WebM recordings, demonstrating capture of both spoken and visual interview behavior\.
#### Role\-conditioned question generation\.
We compare each session’s question set with its own JD and JDs from other position groups\. The median similarity is 0\.055 for the matched JD and 0\.0058 across roles, about a 9\.5\-fold ratio\. The matched JD exceeds the cross\-role baseline in 93\.7% of sessions and ranks first in 82\.4% \(Figure[5](https://arxiv.org/html/2607.10310#S4.F5)\)\. This lexical test does not replace expert judgment, but shows that the generated questions differ systematically by target role\.
#### Adaptive interaction and multimodal assessment coverage\.
The logs contain 1,274 answered main questions, including 208 with follow\-ups and 342 follow\-up turns\. The assessment pool contains 1,425 scored responses because it includes main and follow\-up answers\. These counts and the recordings above show that personalization, multi\-turn interaction, and multimodal assessment operate in the deployed system\.
Figure 5:Workflow coverage, role alignment, and returning\-account evidence\.
## 5Human Expert Evaluation of System Outputs
To evaluate the quality of PolyInterview’s system outputs, we conducted a human expert study\. Ten experts with backgrounds including HR, language assessment, and communication reviewed the same 12 de\-identified, text\-visible artifacts: four role\-conditioned question plans, four response and follow\-up pairs, and four evidence\-linked feedback reports\. They scored three module\-specific criteria from 1 to 5, yielding 120 artifact reviews and 360 criterion ratings\.
Table 3:Scores from the evaluation by ten human experts on a scale from 1 to 5\.
Experts rated question plans most strongly \(4\.62/5\), with high role relevance \(4\.80\) and stage coverage \(4\.83\)\. Follow\-ups were relevant \(4\.65\) but weaker in answer dependence \(3\.28\) and diagnostic depth \(3\.13\)\. Feedback recommendations were actionable \(4\.70\), while polished\-response faithfulness was lower \(2\.75\) because some examples introduced unsupported context\. These results identify strengths and targets for improving follow\-up reasoning and response faithfulness\.
## 6Conclusion
PolyInterview combines questions tailored to the JD and CV with answer\-aware digital human interaction and multimodal assessment grounded in KSA and STAR\. Its hierarchy connects 13 behavior\-level features to 10 assessment aspects and two competency tracks, tracing guidance to multimodal observations\. Across 101 accounts, 1,564 sessions generated 7,665 questions, with 90\.9% of question sets covering all five stages and 82\.4% ranking the matched JD first\. Ten experts found strong question plan quality and actionable recommendations, while identifying follow\-up depth and response faithfulness as improvement targets\.
## Limitations
The usage analysis is observational and the all\-account totals intentionally include internal, test, placeholder, and team activity\. They therefore measure platform activity rather than unique external candidates\. The role\-alignment analysis measures lexical correspondence and still requires expert validation of relevance and quality\. The human\-expert study evaluates representative artifacts rather than longitudinal outcomes\. Because the deployment is recent, we do not yet know whether repeated practice translates into later interview success or employment outcomes\. We therefore plan opt\-in follow\-up surveys with deployed users to track perceived skill gains and subsequent outcomes\. Controlled pre/post studies would still be needed to isolate causal effects\. Artifact\-level expert ratings also do not establish score\-level criterion validity, which requires blinded comparison against independent career\-service ratings\. The interface panels in Figure[3](https://arxiv.org/html/2607.10310#S2.F3)are illustrative walkthroughs from different demo sessions rather than evidence of typical performance\.
## Ethics Statement
PolyInterview is an advisory practice tool for candidates, explicitly not a hiring or gatekeeping system, and its scores are explainable and contestable\. Automated assessment of non\-verbal behavior has a documented history of fairness concerns\. We therefore keep visual assessment advisory, expose supporting evidence for every score, and avoid any pass/fail decision\. The platform processes potentially sensitive data, including CVs, response transcripts, voice recordings, and interview video\. It is deployed through campus servers, with encrypted transmission and storage, strict access controls, and per\-user data isolation\. Raw data are retained only to provide the user\-facing service and are not retained for secondary use\. Users may delete their interview records and associated files at any time, after which they are removed from the server\. Neither PolyInterview nor its external model providers use raw user data to train or fine\-tune models, and provider calls operate under no\-training terms\. Any research use is de\-identified and IRB\-approved, consistent with GDPR and Hong Kong’s PDPO\. The analysis in this paper uses de\-identified session metadata only\.
## References
- Anderson et al\. \(2013\)Keith Anderson, Elisabeth André, Tobias Baur, Sara Bernardini, Mathieu Chollet, Evi Chryssafidou, Ionut Damian, Cathy Ennis, Arjan Egges, Patrick Gebhard, Hazaël Jones, Magalie Ochs, Catherine Pelachaud, Kaśka Porayska\-Pomsta, Paola Rizzo, and Nicolas Sabouret\. 2013\.[The TARDIS framework: Intelligent virtual agents for social coaching in job interviews](https://doi.org/10.1007/978-3-319-03161-3_35)\.In*Advances in Computer Entertainment \(ACE 2013\)*, volume 8253 of*Lecture Notes in Computer Science*, pages 476–491\. Springer\.
- Bai et al\. \(2023\)Jinze Bai, Shuai Bai, Shusheng Yang, Shijie Wang, Sinan Tan, Peng Wang, Junyang Lin, Chang Zhou, and Jingren Zhou\. 2023\.Qwen\-VL: A versatile vision\-language model for understanding, localization, text reading, and beyond\.*arXiv preprint arXiv:2308\.12966*\.
- Baur et al\. \(2013\)Tobias Baur, Ionut Damian, Patrick Gebhard, Kaśka Porayska\-Pomsta, and Elisabeth André\. 2013\.[A job interview simulation: Social cue\-based interaction with a virtual character](https://doi.org/10.1109/SocialCom.2013.39)\.In*Proceedings of the 2013 International Conference on Social Computing \(SocialCom\)*, pages 220–227\. IEEE\.
- Daryanto et al\. \(2025\)Taufiq Daryanto, Xiaohan Ding, Lance T\. Wilhelm, Sophia Stil, Kirk McInnis Knutsen, and Eugenia H\. Rho\. 2025\.[Conversate: Supporting reflective learning in interview practice through interactive simulation and dialogic feedback](https://doi.org/10.1145/3701188)\.*Proceedings of the ACM on Human\-Computer Interaction*, 9\(1\):1–32\.
- Hemamou et al\. \(2019\)Léo Hemamou, Ghazi Felhi, Vincent Vandenbussche, Jean\-Claude Martin, and Chloé Clavel\. 2019\.HireNet: A hierarchical attention model for the automatic analysis of asynchronous video job interviews\.In*Proceedings of the AAAI Conference on Artificial Intelligence*, volume 33, pages 573–581\.
- Hickman et al\. \(2022\)Louis Hickman, Nigel Bosch, Vincent Ng, Rachel Saef, Louis Tay, and Sang Eun Woo\. 2022\.Automated video interview personality assessments: Reliability, validity, and generalizability investigations\.*Journal of Applied Psychology*, 107\(8\):1323–1351\.
- Hoque et al\. \(2013\)Mohammed Ehsan Hoque, Matthieu Courgeon, Jean\-Claude Martin, Bilge Mutlu, and Rosalind W\. Picard\. 2013\.[MACH: My automated conversation coacH](https://doi.org/10.1145/2493432.2493502)\.In*Proceedings of the 2013 ACM International Joint Conference on Pervasive and Ubiquitous Computing \(UbiComp ’13\)*, pages 697–706\. ACM\.
- Inoue et al\. \(2020\)Koji Inoue, Kohei Hara, Divesh Lala, Kenta Yamamoto, Shizuka Nakamura, Katsuya Takanashi, and Tatsuya Kawahara\. 2020\.[Job interviewer android with elaborate follow\-up question generation](https://doi.org/10.1145/3382507.3418839)\.In*Proceedings of the 2020 International Conference on Multimodal Interaction \(ICMI ’20\)*, pages 324–332\. ACM\.
- Janz \(1982\)Tom Janz\. 1982\.Initial comparisons of patterned behavior description interviews versus unstructured interviews\.*Journal of Applied Psychology*, 67\(5\):577–580\.
- Langer et al\. \(2016\)Markus Langer, Cornelius J\. König, Patrick Gebhard, and Elisabeth André\. 2016\.[Dear computer, teach me manners: Testing virtual employment interview training](https://doi.org/10.1111/ijsa.12150)\.*International Journal of Selection and Assessment*, 24\(4\):312–323\.
- Li et al\. \(2023\)Mingzhe Li, Xiuying Chen, Weiheng Liao, Yang Song, Tao Zhang, Dongyan Zhao, and Rui Yan\. 2023\.[EZInterviewer: To improve job interview performance with mock interview generator](https://doi.org/10.1145/3539597.3570476)\.In*Proceedings of the Sixteenth ACM International Conference on Web Search and Data Mining \(WSDM ’23\)*, pages 1102–1110\. ACM\.
- McCrae and Costa \(1987\)Robert R\. McCrae and Paul T\. Costa\. 1987\.Validation of the five\-factor model of personality across instruments and observers\.*Journal of Personality and Social Psychology*, 52\(1\):81–90\.
- Naim et al\. \(2018\)Iftekhar Naim, Md\. Iftekhar Tanveer, Daniel Gildea, and Mohammed Ehsan Hoque\. 2018\.Automated analysis and prediction of job interview performance\.*IEEE Transactions on Affective Computing*, 9\(2\):191–204\.
- Nguyen et al\. \(2025\)Truong Thanh Hung Nguyen, Tran Diem Quynh Nguyen, Hoang Loc Cao, Thi Cam Thanh Tran, Thi Cam Mai Truong, and Hung Cao\. 2025\.SimInterview: Transforming business education through large language model\-based simulated multilingual interview training system\.ArXiv preprint arXiv:2508\.11873\.
- Peterson et al\. \(2001\)Norman G\. Peterson, Michael D\. Mumford, Walter C\. Borman, P\. Richard Jeanneret, Edwin A\. Fleishman, Kerry Y\. Levin, Michael A\. Campion, Melinda S\. Mayfield, Frederick P\. Morgeson, Kenneth Pearlman, Marilyn K\. Gowing, Anita R\. Lancaster, Marilyn B\. Silver, and Donna M\. Dye\. 2001\.[Understanding work using the occupational information network \(O\*NET\): Implications for practice and research](https://doi.org/10.1111/j.1744-6570.2001.tb00100.x)\.*Personnel Psychology*, 54\(2\):451–492\.
- Prajwal et al\. \(2020\)K\. R\. Prajwal, Rudrabha Mukhopadhyay, Vinay P\. Namboodiri, and C\. V\. Jawahar\. 2020\.A lip sync expert is all you need for speech to lip generation in the wild\.In*Proceedings of the 28th ACM International Conference on Multimedia*, pages 484–492\.
- Smith et al\. \(2014\)Matthew J\. Smith, Emily J\. Ginger, Katherine Wright, Michael A\. Wright, Julie Lounds Taylor, Laura Boteler Humm, Dale E\. Olsen, Morris D\. Bell, and Michael F\. Fleming\. 2014\.[Virtual reality job interview training in adults with autism spectrum disorder](https://doi.org/10.1007/s10803-014-2113-y)\.*Journal of Autism and Developmental Disorders*, 44\(10\):2450–2463\.
- Sun et al\. \(2025\)Hongda Sun, Hongzhan Lin, Haiyu Yan, Yang Song, Xin Gao, and Rui Yan\. 2025\.[MockLLM: A multi\-agent behavior collaboration framework for online job seeking and recruiting](https://doi.org/10.1145/3711896.3737051)\.In*Proceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining V\.2 \(KDD ’25\)*, pages 2714–2724\. ACM\.
- Wang et al\. \(2023\)Zihao Wang, Nathan Keyes, Terry Crawford, and Jinho D\. Choi\. 2023\.[InterviewBot: Real\-time end\-to\-end dialogue system for interviewing students for college admission](https://doi.org/10.3390/info14080460)\.*Information*, 14\(8\):460\.
- Zheng et al\. \(2023\)Lianmin Zheng, Wei\-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric P\. Xing, Hao Zhang, Joseph E\. Gonzalez, and Ion Stoica\. 2023\.Judging LLM\-as\-a\-judge with MT\-bench and chatbot arena\.In*Advances in Neural Information Processing Systems \(NeurIPS\) Datasets and Benchmarks Track*\.
## Appendix ADetailed Related Work
#### Commercial interview tools\.
Google Interview Warmup, Yoodli, Big Interview, HireVue, and Interviewing\.io support question rehearsal, selected delivery metrics, response examples, or human coaching\. Newer LLM\-based products additionally offer resume\-tailored sessions and per\-question feedback\. However, as Table[1](https://arxiv.org/html/2607.10310#S1.T1)summarizes, these tools generally do not combine answer\-aware questioning, comprehensive multimodal assessment, and evidence\-linked feedback within one multi\-stage practice workflow\.
#### Virtual interview agents\.
Earlier HCI systems addressed complementary parts of the problem\. MACH provides video\-aligned coaching on smiles, prosody, and filler words\(Hoque et al\.,[2013](https://arxiv.org/html/2607.10310#bib.bib7)\)\. TARDIS/Gloria senses social cues in real time but follows a predetermined scenario\(Anderson et al\.,[2013](https://arxiv.org/html/2607.10310#bib.bib1); Baur et al\.,[2013](https://arxiv.org/html/2607.10310#bib.bib3); Langer et al\.,[2016](https://arxiv.org/html/2607.10310#bib.bib10)\)\. VR\-JIT explains prescripted response options rather than free speech\(Smith et al\.,[2014](https://arxiv.org/html/2607.10310#bib.bib17)\)\. ERICA later generated follow\-ups from answer quality but did not provide post\-interview feedback\(Inoue et al\.,[2020](https://arxiv.org/html/2607.10310#bib.bib8)\)\. Thus, these systems offer multimodal delivery analysis or explainable content feedback, but not both alongside answer\-aware probing\.
#### LLM\-based mock interviews\.
LLMs enable flexible question generation\(Li et al\.,[2023](https://arxiv.org/html/2607.10310#bib.bib11)\), deployed end\-to\-end interviewing\(Wang et al\.,[2023](https://arxiv.org/html/2607.10310#bib.bib19)\), simulations that match people to jobs\(Sun et al\.,[2025](https://arxiv.org/html/2607.10310#bib.bib18)\), talking\-head interaction\(Nguyen et al\.,[2025](https://arxiv.org/html/2607.10310#bib.bib14)\), and transcript\-grounded feedback\(Daryanto et al\.,[2025](https://arxiv.org/html/2607.10310#bib.bib4)\)\. Systems that provide assessment still tend to operate on text alone or cover only selected communication channels\. PolyInterview instead connects adaptive dialogue to a shared rubric spanning textual, vocal, and non\-verbal evidence\.
#### Multimodal and model\-based assessment\.
Automated video interview models predict interview constructs from multimodal features\(Naim et al\.,[2018](https://arxiv.org/html/2607.10310#bib.bib13); Hemamou et al\.,[2019](https://arxiv.org/html/2607.10310#bib.bib5)\), but reliability and validity vary across constructs and individual scores can be difficult to inspect\(Hickman et al\.,[2022](https://arxiv.org/html/2607.10310#bib.bib6)\)\. Foundation\-model evaluators carry additional judgment biases\(Zheng et al\.,[2023](https://arxiv.org/html/2607.10310#bib.bib20)\)\. PolyInterview therefore keeps visual assessment advisory and exposes the intermediate features, aspect mappings, and supporting evidence behind its recommendations\.
## Appendix BTechnology Stack and Service Details
The frontend is a Vue single\-page application, a Flask API orchestrates interview flow and assessment, and a separate Node/Express service handles authentication\. Qwen\-Plus/Qwen\-Flash provide text scoring, Qwen3\-VL\-Plus analyzes non\-verbal behavior\(Bai et al\.,[2023](https://arxiv.org/html/2607.10310#bib.bib2)\), and Qwen3\-ASR supports streaming recognition\. Edge\-TTS provides synthesis and Azure Speech provides pronunciation scoring\. The digital human uses LiveTalking with Wav2Lip\(Prajwal et al\.,[2020](https://arxiv.org/html/2607.10310#bib.bib16)\)and streams over WebRTC\. Four services \(frontend, API, authentication, and digital human\) run behind HTTPS\. Per\-answer WebM recordings are converted to MP4 and WAV for modality\-specific analysis\.
## Appendix CAdaptive Interviewing Details
Personalization operates both before and during the interview\. The planner constructs the initial question sequence from the setup inputs, while the interviewer adapts subsequent dialogue to the candidate’s responses\.
Figure 6:Personalized question generation through description\-to\-aspect and question\-to\-aspect alignment\.#### Question planning\.
The planner receives the target company and position, JD, CV, interviewer persona, and session duration \(Figure[6](https://arxiv.org/html/2607.10310#A3.F6)\)\. It aligns role requirements and question categories with KSA\-based assessment aspects, allocates the duration\-dependent budget across self\-introduction, behavioral, skill\-QA, scenario, and candidate\-question stages, and grounds each question in relevant JD or CV evidence\. This connects personalization to the same competencies used in the subsequent assessment\.
Figure 7:Answer\-aware follow\-up generation using diagnostic scoring, bounded follow\-up rounds, and explicit exit conditions\.
#### Real\-time follow\-ups\.
The interviewer maintains the dialogue context and evaluates each response for contradiction, clarity, and interest \(Figure[7](https://arxiv.org/html/2607.10310#A3.F7)\)\. It can request resolution of a contradiction, clarification of an underspecified response, or elaboration on a point of interest\. The interview proceeds once the response is sufficiently resolved or a section\-time or follow\-up limit is reached\. Basic, Intermediate, and Advanced personas allow at most one, two, and three follow\-ups, respectively, adapting both question content and probing depth\.
## Appendix DAssessment Frameworks: KSA, STAR, and the Big Five
The three frameworks enter the pipeline at different points rather than serving only as background motivation\.
#### KSA: competency specification\.
Knowledge, Skills, and Abilities provides the job\-analysis vocabulary used to organize the cognitive and background aspect families \(Table[4](https://arxiv.org/html/2607.10310#A5.T4)\) and the Professional Competency track\. Aligning planned questions with these aspects connects role requirements to the criteria later used for assessment\.
#### STAR: evidence structure\.
Situation, Task, Action, Result structures the evidence expected in behavioral responses\(Janz,[1982](https://arxiv.org/html/2607.10310#bib.bib9)\)\. PolyInterview uses it to assess whether a candidate explains the context, goal, personal action, and outcome\. A missing element can therefore produce a specific recommendation rather than a generic request for more detail\.
#### Big Five: grounding social constructs\.
The five\-factor model\(McCrae and Costa,[1987](https://arxiv.org/html/2607.10310#bib.bib12)\)informs the selection of communication, interpersonal skills, leadership, and persuasiveness as socially relevant constructs\. PolyInterview does not infer or report personality traits\. The framework only prevents this part of the rubric from relying on ad\-hoc labels\.
## Appendix EMapping Assessment Aspects to Behavior\-Level Features and Question Gating
Each assessment aspect applies a 70/30 weighting to its primary and secondary feature groups \(Table[4](https://arxiv.org/html/2607.10310#A5.T4)\)\. Question gating activates only aspects that a category can reasonably elicit\. Self\-introduction targets Communication Skills, Education & Training, and Experience\. Behavioral questions target Interpersonal Skills, Applied Mental Skills, and Leadership\. Skill\-QA targets Job Knowledge, General Intelligence, and Education & Training\. Scenario questions target Applied Mental Skills, Creativity, and Persuasiveness\. Candidate questions are not scored\. This mapping traces a competency score through its assessment aspects to the underlying textual, vocal, and visual evidence\.
AspectPrimary \(70%\)Secondary \(30%\)*Cognitive*Gen\. IntelligenceConcAcc, LogicCoher, ClarityApplied MentalLogic, ConcAccClarity, CoherCreativityLogic, WordConcAcc, Coher*Background*Job KnowledgeConcAcc, TermLogicEducation & Train\.ConcAcc, TermClarityExperienceConcAcc, Logic, WordTerm*Social*CommunicationClarity, Coher, WordEye, Face, PronInterpersonalEye, Face, CoherPosture, Gest, ClarityLeadershipLogic, Posture, EyeClarity, Face, GestPersuasivenessClarity, Coher, WordEye, Face, ProsodyTable 4:Primary and secondary behavior\-level features used to compute each assessment aspect\.#### Reading the mapping\.
ConcAcc, Term, and Logic denote conceptual accuracy, terminology use, and problem\-solving logic\. Clarity, Coher, and Word denote structural clarity, coherence, and word usage\. Eye, Face, Posture, and Gest are visual features, while Pron and Prosody describe vocal delivery\. Primary and secondary groups indicate relative contribution\.
Gating selects only assessment aspects that a question can elicit\. Inactive aspects remain unscored\. The report traces each competency track through its aspects and behavior\-level features to the supporting evidence, attributing lower scores to specific content, delivery, or non\-verbal signals\.Similar Articles
OmniInteract: Benchmarking Real-World Streaming Interaction for Real-Time Omnimodal Assistants
OmniInteract introduces a streaming benchmark for real-time omnimodal LLMs, evaluating online audio-visual processing with temporal grounding and interactive response requirements. Experiments show that current models perform poorly, with the best overall IA-QTF1 score reaching only 0.368.
PersonaKit (PK): A Plug-and-Play Platform for User Testing Diverse Roles in Full-Duplex Dialogue
PersonaKit is an open-source web platform designed for rapid prototyping and user testing of diverse personas in full-duplex dialogue systems. It allows researchers to configure persona-specific turn-taking behaviors via JSON and conduct A/B surveys to evaluate sociolinguistic interactions.
Human Psychometric Questionnaires Mischaracterize LLM Behavior
This paper finds that human psychometric questionnaires fail to reliably predict LLM behavior in real-world interactions, and proposes generation-based profiling as a more accurate alternative.
SocialPersona: Benchmarking Personalized Profiling and Response with Multimodal Social-Media Context
Introduces SocialPersona, a benchmark for evaluating multimodal large language models on their ability to recover revealed preferences from longitudinal social-media timelines and use them in personalized dialogue.
Examining Human-Like Behaviors in LLMs: A Multi-Dimensional Analysis of Model Behaviors, User Factors, and System Prompts
This paper presents a multi-dimensional analysis of human-like behaviors in LLMs, examining prevalence, effects, and controllability across 21,000 conversations from four models, finding that behaviors vary by model and user factors, with implications for responsible design.