Large language models (LLMs) may support realistic virtual patients, yet most existing systems rely on top-down construction approaches, such as manually curated vignettes. We investigated a bottom-up alternative in which social media histories were transformed into virtual personas examining stability and alignment with human judgment. Histories from Italian Reddit users reporting depressive or anxiety-related distress were manually anonymized and synthesized into two prompt variants (Base and Clinically Enriched). GPT-4o and DeepSeek-V4-Pro completed the PHQ-9 and GAD-7 in independent generations per persona, prompt, model, and scale. Stability was assessed using single-generation and aggregated intraclass correlations. Human alignment with symptom ratings from a psychologist was examined using persona-level correlations and item-level area under the curve (AUC). PHQ-9 total scores showed high single-generation stability across conditions (ICC(2,1) = .84–.88), whereas GAD-7 stability varied by model. Aggregating generations yielded high reliability in all conditions, although item-level stability was heterogeneous. Human alignment was stronger for PHQ-9 than GAD-7 and was highest for DeepSeek-V4-Pro with the Base prompt (r = .98). DeepSeek-V4-Pro also showed higher item-level discrimination, while clinical enrichment had scale-dependent effects. Our findings support the feasibility of bottom-up persona construction, while highlighting the need for psychometric evaluation across models and symptom domains.