Yep, I've definitely seen that. I bet if I were to reserve a portion of the dataset for validation after all tuning, I'd get a much better measure of it's actual ability at prediction. Using it in the way that I did, I wouldn't be surprised if there was significant overfitting. In fact I'd be surprised if there wasn't.